New papers on Efficiency & hardware
86 new papers on efficiency & hardware in the last 7 days, within AI & machine learning. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
Double descent without digital computation
Here we demonstrate double descent in a decentralized analog network of self-adjusting resistive elements.
Peer-reviewed journalClaims a big stepMicrorings as programmable temporal kernels enabling photonic AI beyond 100 Gbaud
Here, we redefine MRRs as programmable temporal convolution kernels by exploiting their impulse responses, enabling computation beyond the resonance linewidth.
Peer-reviewed journalBold claims, read criticallyClaims a big stepReal-world useTinyUDE: Solver-Free Universal Differential Equations on Microcontrollers via Lie-Taylor Jet Matching
We present Lie-Taylor jet matching, a solver-free training framework that fits a hybrid vector field directly to the first and second time-derivatives of observed system states.
PreprintReal-world useFlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention
We present FlashBoB, an exact, I/O-efficient algorithm for BoB in softmax attention that keeps computation within on-chip tiles and avoids all intermediate tensors, where is the sequence length.
PreprintNeural Residual Modeling for Scientific Data Compression under Guaranteed Error Bounds
To the best of our knowledge, this is the first pipeline to combine latent-space RVQ, explicit pixel-space residual correction via a deep spatial post-processing network, and GAE-based error-bound guarantees within a single framework for scientific data compression.
PreprintImplementation of offset corrected AGAD algorithm on 130-nm CMOS technology–based RRAM array for analog neural network training
We present the first implementation of the analog gradient accumulation with dynamic reference (AGAD), reported as the most advanced and highest-performing version of the TT (Tiki-Taka) algorithm, on an HfO 2 -based resistive random-access memory (RRAM) array for analog neural network training.
Peer-reviewed journalBold claims, read criticallyClaims a big stepMemristive singular value decomposition
Here, we present memristive SVD (MSVD), built on compute-in-memory (CIM) memristor chips, enabled by a selective representation enhanced architecture (SREA) that ensures numerical fidelity across iterations.
Peer-reviewed journalBold claims, read criticallyReal-world useTierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching
Across eight text, vision, and audio models on three mobile SoCs, TierKV improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, while incurring only minor accuracy degradation.
PreprintReal-world useNeural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone
We propose Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix.
PreprintBold claims, read criticallyReal-world useInformation Capacity of Generative Video Compression: Quantifying the Rate-Compute Exchange at Identical Quality
To fill this vacancy, we model reconstruction quality as a two-factor power law in data rate and decoder compute, which fits measured DISTS of two GVC decoders with a mean error below 3%, and define the information capacity (IC) as the negative logarithmic slope along an iso-quality contour, namely the fraction of rate saved per fractional increase in compute at identical quality.
PreprintYou Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs
We find that expert selection is far more redundant than the field's operating points assume: uniformly retaining about two thirds of the selected experts preserves 98.8% of unpruned performance on average, requiring only a one-integer change and delivering 1.2-1.7x measured speedup across two serving backends.
PreprintReal-world useTrain Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning
Across four models at 2.79 and 1.88 effective bits, OPD raises average BF16 performance retention from 35% to 70% on MATH-500 and from 66% to 91% on HumanEval while preserving short-form performance, with reasoning gains substantially exceeding those of continued teacher-forced QAD in matched-budget comparisons.
PreprintThe Undetected Damage of Quantization on Retrieval and How to Fix It
We show that a quantized model that keeps its classification accuracy still changes to of its top-1 retrieval results, and that aggregate ranking metrics reveal only part of this damage.
PreprintReal-world useFrom Token Importance to Conditional Removability: Rethinking Visual Token Pruning in Multimodal Large Language Models
Motivated by this perspective, we propose CoRePrune, a training-free two-stage framework.
PreprintIntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
We propose IntBMoE, a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution.
PreprintReal-world useCode availableRisk-Controlled KV-Cache Eviction: From Memory Budgets to Risk Targets
The proposed framework converts a deployment-level reliability requirement into a KV-memory operating point.
PreprintReal-world useExploiting answer-invariant redundancies in satellite imagery for efficient VLM inference on edge
Compared with exhaustive tiled inference, Rift reduces energy by 78% and latency by 69%, while increasing accuracy from 45% to 73%.
PreprintReal-world useFlash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs
In this work, we introduce , a training-free inference acceleration framework for fast and memory-efficient dLLMs.
PreprintReal-world useCode availableSpeech Block Influence: Component-Specific Layer Scoring for Pruning Speech LLMs
We propose Speech Block Influence (SBI), the first layer-importance scoring framework designed for speech LLM pruning that consists of two component-specific scores: SBI-Enc measures the effect of encoder-layer removal at the adapter's output to better reflect downstream impact; SBI-Dec measures layer-wise input-output similarity over text-token positions only to avoid audio-token dominance.
PreprintReal-world useBeyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer Refinement
To address these limitations, we propose Cross-layer Activation-aware Sensitivity Allocation (CASA), a two-phase method.
PreprintAccelerating Video Diffusion via Training-Free Trajectory Routing
We present TRACK: TRajectory-Aware Capacity routing via top-K selection, a heterogeneous denoising strategy that switches between compatible large and small models at selected steps, reducing the average cost per denoising evaluation.
PreprintReal-world useElastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding
We introduce Elastic Threshold Attention (ETA), an end-to-end trainable architecture that achieves hardware-accelerated decoding speed without sacrificing dense model quality.
PreprintBold claims, read criticallyReal-world useBeyond UV Mapping: Mesh Texture Compression via Surface-Aligned Texture Fields
Experiments on the MPEG and AOM mesh compression benchmarks demonstrate improved average rate-distortion performance over representative UV-based methods for both bitstream and GPU-resident compression.
PreprintReal-world useXLOG: A CUDA-Native Engine for Neurosymbolic Integration
xlog is a CUDA-native logic programming engine integrating neural perception with deterministic Datalog, probabilistic inference, and epistemic world views through a typed frontend and provider-owned CUDA runtime.
PreprintFlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates
Building on these observations, we introduce FlashLoop, a training-free inference framework that reduces cross-loop redundancy through token-sparse updates, sparse attention, and KV-residual quantization.
PreprintCompKV: Compensation-Aware KV Selection for Long-Context LLM Inference
To address this limitation, we introduce CompKV, the first compensation-aware sparse attention framework that divides tokens into blocks and explicitly optimizes selection for the downstream compensation mechanism.
PreprintPerplexity Cost Understates What Activation Quantisation Breaks
A perplexity target bounds the average cost of a transformation applied to the activation; it does not, by itself, show which computations survived.
PreprintNGN: Learning Neural Network Size as a Differentiable Count
These results show that structural capacity can be optimized directly as a count.
PreprintDistilling Sequential Computation in Transformer Language Models
We introduce a method for distilling sequential computation by replacing spans of input tokens with collapsed representations, computed on the fly by a lightweight merge module.
PreprintReal-world useBlock-Sparse Attention with Semantic-Geometric Decoupled Routing
To resolve this mismatch, we propose Semantic-Geometric Decoupled Routing, a training-free block routing framework that shifts semantic aggregation to the pre-RoPE space and reconstructs geometric bias with an offline structural prior and relative block distances.
PreprintRGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models
To address these challenges, we propose Riemannian Geometry-Sensitive Quantization (RGSQ), which formulates quantization as a reconstruction problem under a unified Fisher-Riemannian metric.
PreprintCompressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference
We propose a Context-to-Answer-Aligned Memory Compression (CMC) framework, which compresses long input contexts into compact Context Memory Embeddings (CMEs) aligned to any frozen decoder's embedding space, reducing inference costs without modifying decoder weights.
PreprintReal-world useActivation-Flexible ANN-to-SNN Conversion with Finite-State Markov Neurons
We propose a finite-state continuous-time Markov chain (CTMC) neuron framework whose stationary spike flux can approximate every continuous nonnegative monotone activation function on a compact interval.
PreprintStable Geometry with Divergent Task Evidence for Efficient Long-Horizon Agent Compression
Our results show that efficient agent history compression should optimize for preserved task evidence rather than geometric coverage alone.
PreprintPredicting Quantization Price for Selecting PTQ Configurations Before Deployment
We formulate weight-space PTQ as pre-deployment configuration selection using priced layer-output error.
PreprintVGGT-Prime: Compute-Adaptive Mixture-of-Heads for Efficient Visual Geometry Transformers
In light of this observation, we propose VGGT-Prime, a compute-adaptive mixture-of-heads model that resolves this redundancy to accelerate visual geometry transformers while maintaining competitive reconstruction quality.
PreprintDisaggregated Quantization: Specializing LLM Prefill and Decode
We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases.
PreprintReal-world useKV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
We show that co-optimizing rank and bit-width per head, using only standard low-rank projection and scalar quantization, dominates uniform allocation, with the largest gains at low bit-rates.
PreprintReal-world useSLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models
We introduce SLICEChat, a slide-level MLLM that integrates progressive token pruning within a hybrid Mamba--Transformer slide encoder.
PreprintReal-world useCode availableLatent Dataset Distillation for Human Motion Prediction
To address this limitation, we propose a latent DD framework that regularizes distillation with a learned motion prior.
PreprintDiffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding
We therefore introduce GravityOCR, a parameter-shared AR-block-diffusion model jointly trained for parallel drafting and causal AR verification.
PreprintReal-world useRAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models
In contrast, the Jensen-Shannon Divergence achieves zero catastrophic failures, reliably isolating the layers that cannot be safely quantized.
PreprintReal-world useCode availableWhen Quantization Preserves Accuracy but Not Evidence: Explanation-Aware Post-Training Quantization for Medical LLMs
Our objective is to preserve the full-precision model's answer-supporting behavior rather than improve gold-label accuracy, and our method better preserves the full-precision model's answer behavior and rationale-to-answer support.
PreprintReal-world useCode availableMILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression
Experimental results on Qwen2.5 models demonstrate that our method achieves up to 50% reduction in KV cache memory and 1.8x throughput improvement, with negligible performance degradation on classification and reasoning benchmarks, significantly outperforming prior baselines.
PreprintReal-world useLEAP-NBV: Lightweight Edge Active-Perception for Foundation-Model Next-Best-View Planning
We present LEAP-NBV, a lightweight active-perception framework that runs foundation-model-driven Next-Best-View (NBV) planning on-board an edge device.
PreprintReal-world useRBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
We call this failure mode mean dilution and propose RBS-Attention, a training-free sparse-prefill method with two complementary selection branches.
PreprintReal-world useArtificial Structure Function Search: Preserving Artificial Functional Connectivity for Structured Pruning
We present results for PGI as a selection criterion and for ASF-S as a pruning framework against recent benchmarks, demonstrating that our method yields model variants with 70% parameter reduction, that can recover baseline accuracy without re-training the pruned layers.
PreprintBold claims, read criticallyReal-world useAccelerating Dense LLMs via L0-regularized Mixture-of-Experts
In this paper, we propose L0-MoE, a lightweight MoE approach using L0-regularization to accelerate dense LLMs nearly without performance loss.
Preprint with a published versionReal-world useSix Layers Less: Encoder Pruning for Whisper with Label-Free Recovery
We present an approach that ranks encoder layers by the leave-one-layer-out change in Word Error Rate (WER).
PreprintReal-world useCode availableYou've Seen Enough: Quality-Constrained Image Coding for Machines
Experimental results show that, under the quality constraint, the proposed method achieves a BD-rate of over an unconstrained joint rate--distortion--task optimization and over a simple rate--distortion baseline, showcasing bitrate reduction with the same task performance.
Preprint