Foundations
Mathematical Foundations
Linear algebra, calculus, probability, statistics and optimisation — the maths behind ML
Skills in this topic
4 skills — Sign in to track your progress
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1w ago
EVINCE: Optimizing Multi-LLM Dialogues Using Conditional Statistics and Information Theory
arXiv:2408.14575v5 Announce Type: replace Abstract: EVINCE (Entropy and Variation IN Conditional Exchanges) is a novel framework for optimizing multi-LLM dialog
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1w ago
Gap Entropy and Almost Instance-Wise Optimal Best-Arm Identification
arXiv:2609.13703v1 Announce Type: cross Abstract: In the best-arm identification problem, we are given $n$ stochastic arms with unknown means and wish to identi
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1w ago
Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression
arXiv:2609.15838v1 Announce Type: cross Abstract: Per-matrix singular value decomposition (SVD) truncation is Eckart-Young optimal in the whitened Frobenius nor
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1w ago
Constrained Online Learning with Noisy Constraint Values
arXiv:2609.06921v2 Announce Type: replace-cross Abstract: We study constrained online convex optimization with adversarial constraints and conditionally unbiase
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1w ago
Linear Exponential Quadratic Gaussian Covariance Steering
arXiv:2609.12463v1 Announce Type: cross Abstract: We formulate and analyze the linear exponential quadratic Gaussian (LEQG) covariance steering problem in conti
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
3w ago
Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs
arXiv:2608.28853v1 Announce Type: cross Abstract: Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-ord
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
3w ago
Denoising as Projection: Constrained Optimization with Gradient-Guided Diffusion
arXiv:2608.29507v1 Announce Type: cross Abstract: Diffusion models are increasingly used not only for sampling from learned data distributions, but also for gen
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
3w ago
EquiReg: Equivariance Regularized Diffusion for Inverse Problems
arXiv:2505.22973v3 Announce Type: replace-cross Abstract: Diffusion models represent the state-of-the-art for solving inverse problems such as image restoration
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
3w ago
KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution Preservation
arXiv:2608.27839v1 Announce Type: new Abstract: Fine-tuning-based knowledge editing is simple and architecture-agnostic, but standard cross-entropy increases th
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
3w ago
Marginal Coverage Credit Reduces Redundant Exploration in Parallel State-Entropy Optimization
arXiv:2608.27507v1 Announce Type: cross Abstract: Policy Gradient for Parallel State Entropy maximization (PGPSE) expands state-space coverage by training indep
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
3w ago
Physics-Informed Stochastic Configuration Machine: A Backpropagation-Free Neural Network with Fast Training for Nonlinear Differential Equations
arXiv:2608.26549v1 Announce Type: cross Abstract: While Physics-Informed Neural Networks (PINNs) have emerged as a transformative paradigm for solving complex d
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
4w ago
Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning
arXiv:2608.24386v1 Announce Type: cross Abstract: Tensor-valued prediction is fundamental to geometric deep learning, yet uncertainty quantification (UQ) for su
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
4w ago
Across the Loss Landscape with Progressive Growth
arXiv:2608.24568v1 Announce Type: cross Abstract: Deep neural networks generalize well despite their highly nonconvex, overparameterized loss landscapes, a phen
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
4w ago
A Physical Response-and-Memory Model for Muon Optimization
arXiv:2608.22994v1 Announce Type: cross Abstract: Training large language models is costly. How low a loss the same compute can ultimately reach depends on how
MarkTechPost
🔢 Mathematical Foundations
4w ago
Google Research Introduces ME-POIs: A Mobility-Informed Framework that Adds “How a Place Is Used” to Text-Based POI Embeddings
framework that folds aggregate human movement into text-based place embeddings. Language models describe what a place is; they miss how it is used. ME-POIs enco
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers
arXiv:2608.20802v1 Announce Type: new Abstract: Human motion forecasters are increasingly accurate and fast, but reliable deployment requires uncertainty estima
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Graphon Particle Systems, Part II: Dynamics of Distributed Stochastic Continuum Optimization
arXiv:2407.02765v4 Announce Type: replace-cross Abstract: We study the distributed optimization problem over a graphon with a continuum of nodes, which is regar
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Entropy-Constrained Adaptive Stochastic Quantization
arXiv:2608.18147v1 Announce Type: cross Abstract: Adaptive stochastic quantization (ASQ) is a recently introduced quantization approach that optimizes the Mean
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Automated Computational Energy Minimization of ML Algorithms using Constrained Bayesian Optimization
arXiv:2407.05788v2 Announce Type: replace-cross Abstract: Bayesian optimization (BO) is an efficient framework for optimization of black-box objectives when fun
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression
arXiv:2608.17434v1 Announce Type: new Abstract: We study Gaussian regression over the explicit vector-valued Parhi--Nowak deep-RBV^2 architecture with depth L,
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization
arXiv:2502.00213v5 Announce Type: replace-cross Abstract: Transformers are difficult to optimize with stochastic gradient descent (SGD) and largely rely on adap
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Solving nonconvex Hamilton--Jacobi--Isaacs equations with PINN-based policy iteration
arXiv:2507.15455v3 Announce Type: replace-cross Abstract: We propose a mesh-free policy iteration framework that combines classical dynamic programming with phy
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL
arXiv:2608.14559v1 Announce Type: new Abstract: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} t
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
arXiv:2608.16697v1 Announce Type: new Abstract: Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation in
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data
arXiv:2608.14712v1 Announce Type: cross Abstract: Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning
arXiv:2608.13621v1 Announce Type: new Abstract: A hidden Markov model (HMM) combines three roles: inference of a hidden-state belief from observations, propagat
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
BGA: A noise-immune neural distillation framework for malicious signature extraction in high-entropy encrypted flows
arXiv:2608.14126v1 Announce Type: cross Abstract: To mitigate attention dilution in high-entropy TLS 1.3 flows, we propose BGA, a noise-immune neural distillati
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Multi-Objective Bayesian Optimization for Model Merging
arXiv:2608.14264v1 Announce Type: cross Abstract: Model merging combines trained models directly in weight space, offering a compute-efficient alternative to ad
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning
arXiv:2608.11749v1 Announce Type: cross Abstract: Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating t
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation
arXiv:2608.12196v1 Announce Type: cross Abstract: Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-drive
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Control-Oriented Scenario Tree Construction through Reinforcement Learning
arXiv:2608.09335v1 Announce Type: new Abstract: Multistage stochastic model predictive control (MPC) handles uncertainty by optimizing over a scenario tree, a f
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Two-Step MV-DeepONet: Probabilistic Operator Learning for Uncertainty Propagation Driven by Random Input Fields
arXiv:2608.09071v1 Announce Type: cross Abstract: Forward uncertainty propagation in complex physical systems can induce structured covariance across field-valu
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
1mo ago
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
arXiv:2307.10053v5 Announce Type: replace-cross Abstract: In this paper, we focus on providing convergence guarantees for stochastic subgradient methods in mini
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
⚡ AI Lesson
2mo ago
Tangent classes of matroids and wonderful compactifications
arXiv:2607.05835v1 Announce Type: cross Abstract: For every loopless matroid $M$ and every Feichtner--Yuzvinsky building set $\mathcal{G}$ containing the top fl

MIT Technology Review
🔢 Mathematical Foundations
⚡ AI Lesson
3mo ago
Super Mario is mathier than you think
Here’s a problem you probably didn’t solve in school: You’re an ambitious young plumber from Brooklyn in a world inhabited by violent human-size mushrooms calle
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
3mo ago
Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling
arXiv:2606.09926v1 Announce Type: cross Abstract: Sampling from the sequence-level power distribution $p^\alpha$ elicits RL-level reasoning from base language m
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
3mo ago
Mind the Gap: Mixtures of Gaussians in Approximate Differential Privacy
arXiv:2605.28078v1 Announce Type: cross Abstract: We design a class of additive noise mechanisms that satisfy \((\varepsilon, \delta)\)-differential privacy (DP
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
⚡ AI Lesson
4mo ago
Amortized Energy-Based Bayesian Inference
arXiv:2605.15407v1 Announce Type: cross Abstract: We consider amortized Bayesian inference for nonlinear inverse problems in settings where only samples from th
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
⚡ AI Lesson
4mo ago
$f$-Trajectory Balance: A Loss Family for Tuning GFlowNets, Generative Models, and LLMs with Off- and On-Policy Data
arXiv:2605.15417v1 Announce Type: cross Abstract: In GFlowNets and variational inference, it has been shown that the mean square error between target and model
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
⚡ AI Lesson
4mo ago
PrismQuant: Rate-Distortion-Optimal Vector Quantization for Gaussian-Mixture Sources
arXiv:2605.15507v1 Announce Type: cross Abstract: For a Gaussian source under mean-squared error (MSE), classical transform coding is rate--distortion (RD) opti
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
4mo ago
Fast Rates for Inverse Reinforcement Learning
arXiv:2605.14599v1 Announce Type: cross Abstract: We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement le
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
⚡ AI Lesson
4mo ago
Deterministic Decomposition of Stochastic Generative Dynamics
arXiv:2605.08794v1 Announce Type: cross Abstract: Modern generative models can be understood as probability transport from a simple base distribution to a targe
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
⚡ AI Lesson
4mo ago
Every finite group admits a just finite presentation
arXiv:2605.10402v1 Announce Type: cross Abstract: A finite presentation of a finite group is called `just finite' if removing any relation from R results in a p
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
⚡ AI Lesson
4mo ago
A Milestone in Formalization: The Sphere Packing Problem in Dimension 8
arXiv:2604.23468v2 Announce Type: cross Abstract: In 2016, Viazovska famously solved the sphere packing problem in dimension $8$, using modular forms to constru
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
⚡ AI Lesson
5mo ago
HHL with a Coherent Fourier Oracle: A Proof-of-Concept Quantum Architecture for Joint Melody-Harmony Generation
arXiv:2604.20882v1 Announce Type: cross Abstract: Quantum algorithms with a proven theoretical speedup over classical computation are rare. Among the most promi
ArXiv cs.AI
🔢 Mathematical Foundations
📄 Paper
5mo ago
A Nonasymptotic Theory of Gain-Dependent Error Dynamics in Behavior Cloning
arXiv:2604.14484v1 Announce Type: cross Abstract: Behavior cloning (BC) policies on position-controlled robots inherit the closed-loop response of the underlyin
DeepCamp AI