New papers on Learning theory & optimization
113 new papers on learning theory & optimization in the last 7 days, within AI & machine learning. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning
To the best of our knowledge, this is the first work to quantify cross-task complexity for general nonlinear task families and explicitly construct a Transformer that exploits their low-dimensional structure to perform ICL.
PreprintClaims a big stepContinuous Optimization for p-adic Models
We present the first method for native, continuous gradient descent for machine learning models with -adic parameters.
PreprintBold claims, read criticallyClaims a big stepCode availableToward a Unified Mathematics of Concepts
We propose an operation-based view that evaluates mathematical frameworks by the conceptual operations they support, identifying thirteen operations (including similarity, composition, generalization, and grounding) that recur across cognition, psychology, and AI.
PreprintDisplacement Geometry Captures Platonic Shared Reality Across Models and Modalities
In this paper, we show that what models share is not the location of samples in representation space, but the directions (displacement vectors) between them.
PreprintBold claims, read criticallyReal-world useAdversarially Robust PAC Learning with Optimal VC Rates
We determine the optimal -independent sample complexity of this problem in both the realizable and agnostic settings.
PreprintxWhyL: Causal Interactive Learning
To fill this gap, we propose xWhyL, a formal framework connecting causality and XAI by learning causal models from explanations.
PreprintScalable Minimum-Volume Simplex Estimation with Non-asymptotic Analysis
We propose DeepMVSA, which re-expresses the minimum-volume principle in neural implicit form: a lightweight coordinate network generates the mixing weights and a triangular LU-type parameterization the dual simplex matrix, reducing the trainable-state memory to , independent of , and the cost per data pass to .
PreprintNuclear Norm-Regularized Bayesian Matrix Completion
We give the first sampler for this model with an explicit non-asymptotic guarantee: polynomial in the matrix dimensions and in the reciprocal of the target accuracy.
PreprintGap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy
We give a new analysis of the ubiquitous Oja's algorithm [Oja82] for the most general, gap-free variant of this problem, where no eigengap assumptions are made on the underlying mean matrix, complemented by a nearly-matching lower bound.
PreprintHuman-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation
SCALE is valid at finite sample sizes and matches the lower bound to first order as the target error probabilities vanish.
PreprintNeutral-Atom-based Quantum Optimization for Resource Allocation in NOMA Networks
In this paper, we investigate the use of neutral-atom quantum platforms to solve the maximum access problem (MAP), formulated as a mixed-integer programming task that jointly considers admission control, user clustering, channel assignment, and power allocation in a non-orthogonal multiple access (NOMA)-enabled uplink network.
PreprintReal-world useThe Capability Manifold and ML Scaling Laws
We bridge this gap by introducing a capability manifold, a multidimensional framework mapping downstream capabilities to pre-training, post-training, and test-time resources through bounded scaling functions.
PreprintConformal Robustness in Prediction-Driven Decision-Making
Overall, our work shows that conformal scores endow fixed black-box predictors with an interpretable uncertainty scale for downstream decision-making while enabling reliability guarantees, acceptable-target selection, and fragility analysis within a unified framework.
PreprintReal-world useDirect Message Approximation (DMA): A Consistency-Based Framework for Tractable Approximate Inference on Factor Graphs
We introduce Direct Message Approximation (DMA), which approximates factor-to-variable messages directly rather than the marginal.
PreprintStochastic Inertial Krasnosel'skii-Mann Iteration Achieves Near-Optimal Sample Complexity
To our knowledge, this is the first single-loop method for general nonexpansive fixed-point problems to attain this near-optimal sample complexity without variance reduction or batching.
PreprintBilevel Optimization of Topology and Hyperparameters (BOTH)
Here, we propose differentiating TO itself using automatic differentiation.
PreprintBold claims, read criticallyWhen Does Unsupervised Learning Succeed or Fail? A PoS Perspective on Reconstruction-Based Anomaly Detection
These results connect reconstruction failures to identifiable geometric conditions and provide practical mechanisms for learning compact representations.
PreprintReal-world useResource-Adaptive Stochastic Gradient Descent for Online Linear Programming without Re-solving
Under standard non-degeneracy conditions, our algorithm is feasible on every sample path and achieves O(\log T) expected regret against the realized fractional hindsight optimum, which matches the lower bound, even for policies that know the distribution and have unrestricted computation.
PreprintReal-world useDouble Descent and Malign Overfitting in Diffusion Models
This overfitting is malign because, although the implicit regularization of training is fully at work, it drives the model toward the empirical score, which memorizes the training set, rather than toward the true score.
PreprintWhitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures
These results indicate that the squared whitened norm is better interpreted as a Mahalanobis measure of semantic atypicality than as a log-likelihood.
PreprintA Smoothed Discrepancy Principle for Random Feature Methods and Neural Networks
Building on the discrepancy principle, we propose a multi-scale stopping rule that applies to general kernel estimators and show that, unlike previous approaches, it achieves full adaptivity over all smoothness levels in the well-specified case.
PreprintA Spectral Theory of Grokking: Weight Decay induces Feature Learning
We provide a quantitative theory for how this transition from lazy to rich learning can produce delayed generalization.
PreprintThe Price of Self-Calibration: Exact Evidence Budgets and Manufactured Blind Sets in Adaptive Monitoring
Third, any monitor required to tolerate a drift class is blind, at any horizon and for any rule, to every fault in ; the proof is a deliberately elementary two-point argument and the contribution is the object it identifies: for speed-bounded classes the blind set is exactly the doubled-speed class, and the tracker absorbs a speed class fixed by its own gain, so that under a certification regime declaring absorbed drift normal, the monitor manufactures .
PreprintPROSE: A Theory of Optimal Stopping with Perishable Evidence for Peer Selection in Intermittently Connected Decentralised Learning
This paper develops a self-contained theory of optimal stopping for the resulting peer-selection problem.
PreprintExtreme classification: beating chance with one training example from each class
A measurable encoding then gives a deterministic distribution-free rule which beats chance on every countably generated measurable space, in particular every separable metric space.
PreprintOrbital Error Dynamics: Self-Organized Criticality, Ephemeral Parameter Resonance, and Non-Linear Biological Ontologies in Zero-Storage Neural Synthesis
We further couple an enteric-cranial Dual-Brain architecture shielded by adaptive CD4+ regulatory immune gating (M_CD4), and project the 4-nucleotide genetic basis (A, T, C, G) across quadrants in C. Multi-seed empirical validation on the Two-Moons manifold (5 seeds, 80/20 train/test split, 32x32 grid, zero test-time updates, zero label leakage) demonstrates that procedural parameterization from a 24-byte coordinate seed achieves 77.67% +/- 5.35% clean test accuracy (within an 8.00-point paired difference of an unconstrained gradient baseline at 85.67% +/- 5.35%, 95% CI: [-1.07%, 17.07%]) and 71.33% +/- 3.80% under distribution shift (N(1.2, 0.4)), alongside conceptual equivalence with an analog optical co-processor.
Preprint with a published versionBold claims, read criticallyCode availableWhy Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning
Our unified analytical framework provides a rigorous mathematical explanation for three central empirical puzzles in SL: (i) under shared initialization, the transfer operator forms a strictly Positive Semi-Definite (PSD) structure, guaranteeing that ghost-output optimization aligns the student with the teacher's true task objective without explicit label exposure; (ii) the ghost-output dimensionality acts as an explicit rank bottleneck governing the transfer of task-relevant features; and (iii) synthetic, high-entropy inputs function as broadband probes that maximize cross-task kernel overlap, explaining why random noise consistently outperforms structured data for subliminal transfer.
PreprintBold claims, read criticallyDiagonalized Attention for Individualized Regression: Latent-Row Localization and Prediction
To estimate this model, we introduce a diagonalized attention mechanism that uses query--key scores to localize sample-specific signal rows and a value matrix for downstream regression.
PreprintSuperPCA: subspace analysis and an efficient algorithm for high-dimensional PCA
Our main theoretical finding is that the subspace spanned by several leading eigenvectors of the sample covariance matrix contains significant information about the desired signals long before the individual eigenvectors converge to the population principal components.
PreprintStatistical Gains from Looped Estimation under Parameter Budgets
For targets of known H\"older smoothness, looped residual feedforward networks and a specified post-layer-normalized Transformer attain the minimax polynomial rate up to logarithmic factors with a fixed number of bounded real parameters.
PreprintMF-SCBO : Multi-fidelity Scalable Constrained Bayesian Optimization
In this work, we extend the Scalable Constrained Bayesian Optimization method to the multi-fidelity setting, resulting in the MF-SCBO method.
PreprintWhat Must Survive? Exact Task-Information--State Frontiers for Resource-Sufficient Learning
For an advice alphabet of size , the exact frontier is \[ p^*(K)= \min_{\substack{\Pcal partition of \U\\|\Pcal|\le K}} \max_{C\in\Pcal}\rank(T_C), \] with the -bit frontier obtained by setting .
PreprintSelf-Supervised Combinatorial Optimization with Constraints via Frank-Wolfe
We propose a general framework in which the neural network is allowed to predict arbitrary continuous vectors that could potentially lie outside of the feasible polytope.
PreprintFull-Covariance Smoothing of Bayesian Neural Networks for Online Adaptation
We overcome this limitation via a cross-covariance identity that enables full-covariance propagation through a network's nonlinear activations.
PreprintOn the SoS Certifiability of Log-Concave Distributions
For an arbitrary isotropic log-concave distribution on , we prove that the polynomial is a sum of squares for every even , where is a universal constant.
PreprintOn the Sample Complexity of Active Learning with Membership Queries
In particular, some hypothesis classes that are inherently slow to learn in the pool-based setting, achieving only polynomial error decay in the number of samples, become exponentially learnable once synthesized queries are allowed.
PreprintOptimal Randomized Proper Online Learning
We prove that the optimal expected mistake bound of online learning a function class by a randomized proper learning algorithm is , where is the Littlestone dimension of and is the time horizon.
PreprintSelective Inference for Deep Clustering in Latent Spaces
In this work, we develop an SI framework for deep clustering with a fixed pretrained encoder.
PreprintSequential Confidence Sets for Coverage-Constrained Conformal Model Selection
We introduce Coverage-Constrained Sequential Model Confidence Sets (CC-SMCS), which separate certifiably feasible, possibly feasible, and possibly constrained-optimal pipelines.
PreprintRACER: Role-Aligned Competence Estimation for Human-AI Routing
We propose RACER---Role-Aligned Competence Estimation for Routing---a role-relative framework for estimating an unseen expert's competence from context.
PreprintReal-world useThe Ups and Downs of Backprop Weights
We propose weight operators: parameterized modules that implement reusable functional components and can be composed at inference to form the function required by each sample.
PreprintMatrix Aggregation Operators
This paper addresses this gap by formalizing the notion of a matrix aggregation operator (MAO).
PreprintIn-Context Guidance: Learning Inter-Task Synergies via Numerical Foundational Models for Few-Shot Multitask Optimization
This paper introduces In-Context Guidance Multitask Optimization (ICG-MTO), a novel framework that leverages numerical foundational models to improve inter-task coupling estimation in few-shot scenarios.
PreprintSparse Priors for Efficient Distribution Learning
We show that distribution learning under a -sparse prior achieves a Bayesian risk lower bound of under common distance metrics, and show a matching (up to logarithmic terms asymptotically in ) upper bound for the TV distance under mild additional assumptions.
PreprintWhat Converges in the Platonic Representation Hypothesis? Structure over Geometry
Together, these results show that relational convergence extends beyond local neighborhoods to global spanning structure, whereas metric geometry exhibits substantially weaker convergence.
PreprintPAC-Bayesian Meta-Learning for Few-Shot Identification of Linear Dynamical Systems
We propose PBML-LTI, a PAC-Bayesian meta-learning framework for few-shot LTI system identification that learns a transferable prior over task-specific dynamics while preserving task heterogeneity.
Preprint with a published versionSensing to Intelligence: Principles for Neuromorphic Circuits and Systems
A system is truly neuromorphic when the invoked biological or physical principle plays a causal, design-relevant role in how information is represented, how state evolves, or what capability the complete system achieves.
PreprintAn Analytical Theory of Auxiliary Learning
For non-linear activation functions, we develop a fluctuation-dissipation analytical theory that establishes a general relation linking the main and auxiliary errors to the corresponding single-task error.
PreprintTracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking
Our results connect partial accuracy, learning stages, and internal computation through the subgroup cosets that models learn to track.
PreprintThe Dynamics of Quasiregular Neural Learning
Our results isolate a simple form of competition between regularities and exceptions during neural learning.
Preprint