New papers on Reinforcement learning
106 new papers on reinforcement learning in the last 7 days, within AI & machine learning. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input.
PreprintWhich Constraints Are Missing? Ask the Verifier: Graded Rewards for Constraint-Following Music Generation
We therefore introduce MusicRLVR, which pays graded per-property credit behind a hard validation gate that rejects malformed outputs, plus a joint-satisfaction bonus, requiring no human annotation, learned reward model, or music-domain fine-tuning.
PreprintProvably Efficient Reinforcement Learning in Continuous-Time Episodic MDPs with Poisson Decision Epochs
These results provide the first tight regret guarantees for Lipschitz-smooth continuous-time episodic MDPs with Poisson decision epochs.
Preprint with a published versionVISTA: An Attention-Based Multi-Agent Reinforcement Learning Architecture for Space Situational Awareness Sensor Tasking
We introduce VISTA (Variable-Entity Intelligent Sensor Tasking Architecture), a scalable deep reinforcement learning architecture for persistent uncertainty-driven catalogue maintenance across variable object populations and sensor configurations.
PreprintReal-world useCode availableThe Evidence Ladder for Reinforcement Learning in Healthcare: From Retrospective Policies to Trusted Interventions
The central lesson is that healthcare RL should be evaluated as an intervention embedded in a changing sociotechnical system, rather than only as an optimizer of a retrospective reward.
PreprintOneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios
We present OneBid, a unified auto-bidding foundation model that learns a reusable backbone from heterogeneous oCPX logs and adapts it to scenario-specific deployments via offline post-training.
PreprintReal-world useFully Byzantine-Resilient Multi-Agent Reinforcement Learning
Under linear parameterizations of the value and team-reward functions and Byzantine edge attacks, where adversarial behavior is confined to the communication layer, we prove that agents' parameters converge almost surely to the same limit points as in the attack-free case over time-varying communication graphs.
PreprintAutoGym: Blueprint-First Generation of Verifiable Agent Gyms
We present AutoGym, a framework that generates complete gyms (tasks, executable environments, and verifiers) from a minimal domain seed or prior model trajectories.
PreprintProcessLight: Process Supervision for Large Language Model Based Traffic Signal Control
To bridge this gap, we propose an LLM-based framework ProcessLight to decompose signal decisions into verifiable semantic steps.
PreprintReal-world useCode availablePolicy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning
We derive MI-SARSA, an on-policy temporal-difference algorithm that incorporates mutual-information regularization through a learned marginal action prior and a penalty on state-specific deviations from that prior.
PreprintAdversarial Closed-Loop Curriculum for Evolving Role-Playing Agents
Therefore, we propose AdvRole, an adversarial context rewriting framework that turns role-playing RL into a closed-loop curriculum.
PreprintA General Framework for Budgeted Threshold Incentives on Request
We present a request-driven framework that composes four stages (conditional prediction, population reduction, trajectory integration and budget allocation) through seven replaceable modules that exchange conditional trajectory laws, whose award probabilities and award-marked moments give payment and uplift for any activity rule.
PreprintReal-world useRULER: Instance-aware Rubric Rewards for SVG Generation
Building on this finding, we introduce RULER (Instance-aware Rubric Rewards for Reinforcement Learning), which converts each instruction into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes; a judge VLM scores rendered rollouts item-by-item, and the weighted satisfactions form a fine-grained reward optimized via Group Relative Policy Optimization.
PreprintAutonomous Model Lifecycle Management for Digital Twin-Based Manufacturing Control
This paper presents a closed-loop Cyber-Physical System (CPS) for autonomous model lifecycle management in automotive manufacturing, deployed since 2023.
PreprintReal-world useRLVR landscapes for iterated multiplications can be benign: Insights from spin-glass theory
We show, both theoretically and experimentally, that for a wide class of models and tasks with uncorrelated inputs, this landscape is benign, containing no local minima that could trap RLVR training.
PreprintGraph-Based Inference and Topology-Aware Multi-Agent Reinforcement Learning for Large-Scale Railway Network Management
A central contribution of this work is the demonstration of scalability through zero-shot transfer learning: graph-based agents, trained only on small network portions, are successfully deployed in a zero-shot manner on large-scale unseen networks without any retraining.
PreprintBold claims, read criticallyReal-world useThe Right Future for Action: Learning Action-Relevant Predictive States in World Action Models
These findings motivate Action-Relevant Predictive States (ARPS), a compact predictive interface between the video and action experts.
PreprintError- and Prediction-Driven Motor Learning in the Cortico-Cerebellar Loop
Our results show that single-signal predictions are insufficient under delay, while multiplexing and feedback together provide a unified mechanism for online control and rapid learning.
PreprintPoEM: Predicting RL Outcomes from Existing Policies
We answer this in the affirmative by introducing PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards.
PreprintPACT: From Credit Assignment to Critic Alignment
We formulate three regularity conditions, namely Completeness, Prefix Consistency, and Neutrality, and prove that they uniquely determine token-level credit.
PreprintLearning Defensive Policies against Diverse Inference Attacks for Smart Meter Privacy
We propose a proxy-guided hierarchical reinforcement learning framework that learns battery-based load-shaping policies to inject realistic but misleading appliance-level signatures into the aggregate signal, thereby disrupting the structured patterns exploited by NILM.
PreprintReal-world usePCQC: Privileged Counterfactual Question Credit for Multi-Turn Medical Dialogue
To address this gap, we introduce PCQC (Privileged Counterfactual Question Credit), which uses privileged patient information during training to learn from questions never asked.
PreprintReal-world useSynthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search
We present a method for synthesizing reactive character behaviors for continuous games as compact, human-readable programs.
Preprint with a published versionRLVR: Reinforcement Learning with Verifiable Rubric-based Ranking
We propose Reinforcement Learning with Verifiable Rubric-based Ranking (RLVR), a verifiable ranking paradigm for rubric-based RLVR.
PreprintRLVR is a Kernel, Not a Function: Statistical Inference for pass@ Crossovers
The relationship is a conditional distribution---a Markov kernel---rather than a single curve, and fitting it predicts crossings in independent generations for the same prompts and corrects the simple model's power estimates.
PreprintCanopy: Exploiting Piecewise Smooth Tree Priors for Multi-Fidelity Bandits
We introduce CANOPY, a multi-fidelity tree bandit that learns where the smoothness prior is valid rather than assuming it globally.
PreprintReal-world useArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL
We propose ArenaFlow, a hierarchical credit propagation framework for open-ended agent reinforcement learning.
PreprintMinimal Recurrent Behavioral Memory for Imitation under Partial Observability
We characterize this minimal recurrent behavioral memory by a compatibility relation: under transitivity its classes attain the exact minimum, while the general case is an entropy minimization over closed compatible state assignments, with exact certificates on finite instances.
PreprintCode availableInformation-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation
We introduce CoVer (Co-trained Coder and Verifier), a single-policy GRPO framework that addresses both failure modes.
PreprintOnlineWM: Causality-Aware Active Online Learning for Effective World Modeling
To address these limitations, we propose OnlineWM, an online training framework that continuously improves world modeling through active simulator interaction and causality-aware optimization.
PreprintContext-Continuous Preference Learning for Exoskeleton Personalization
We propose Context-Continuous Preference Learning (CCPL), a Gaussian-process preference model that shares observations across nearby contexts while retaining context-specific utility estimates.
PreprintReal-world useEBRL: Asynchronous Embodied RL by Multi-Grained Resource Management
Experiments show that EBRL achieves 1.30-3.47 times the end-to-end rollout throughput and 2.5 times of training convergency compared to the SOTA embodied RL systems.
PreprintOptimal No-Regret Learning for Repeated Prophet Inequality
We give an efficient algorithm achieving expected regret, matching the lower bound up to logarithmic factors.
PreprintDeep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing
In this paper, we present a novel end-to-end, size-agnostic graph reinforcement learning framework for 1D-BPP.
PreprintSearch-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search
We present a search-aware reinforcement learning (RL) framework for QU based on a distill-then-RL paradigm.
PreprintReal-world useDCRL: Decoupling and Coupling Reinforcement Learning via Policy-Reward Manifold Alignment
To address the aforementioned misalignment, we propose Decoupling and Coupling Reinforcement Learning (DCRL) framework, which incorporates two key components: (1) a syllogistic logic-based prompt evolution mechanism that dynamically refines reward rubrics to enhance the expressiveness of the reward manifold; and (2) a policy-reward re-coupling mechanism that jointly updates the reward and policy models, ensuring consistent evaluation and mitigating manifold mismatch during training.
PreprintBold claims, read criticallyEAGER: Enhancing Generative Event Extraction via Reinforcement Learning with Verifiable Rewards
We present EAGER, a reinforcement learning framework for generative event extraction that combines fine-grained verifiable rewards with Schema-Contrastive Advantage Estimation to alleviate advantage collapse under sparse binary rewards.
PreprintLearning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks
This paper presents a shared latent-space framework that connects simulator calibration and reinforcement learning control through a common learned representation of urban traffic dynamics.
PreprintBold claims, read criticallyReal-world useOn Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning
We show that this extension is frequently harmful: across four preference-conditioned off-policy algorithms spanning two critic backbones and two preference-sampling schemes on the continuous-control MO-Gymnasium suite, it degrades 19 of 36 algorithm-environment settings by as much as four standard deviations, improves only one, and leaves the rest unaffected.
PreprintCritical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use
We introduce Critical-State RL to identify trainable states in multi-turn interactions.
PreprintMATES: Learning Multi-Agent Interactions by Transforming Observations for Frozen Single-Agent Policies
We introduce Multi-Agent Observation Transformation for Existing Single-Agent Policies (MATES), an input-side adaptation framework for tasks whose multi-agent observations preserve the solo-task information while exposing separately identifiable neighbor information.
PreprintInformation-Time Proximal Policy Optimization
In this paper, we propose InfoPPO, which reparameterizes temporal progression using information density rather than raw token count.
PreprintTowards Full Pipeline FP8 Reinforcement Learning for LLMs
Extensive experiments across GRPO and DAPO algorithms, model scales from 8B to 32B, and multiple FP8 scaling granularities demonstrate that our approach successfully eliminates entropy surges and restores performance comparable to the BF16 baseline.
PreprintExact Bayes Regret and Asymptotic Optimality in High-Dimensional Gaussian Bandits
A policy-uniform lower bound identifies the limiting optimal Bayes regret and proves that posterior-mean greedy selection attains it.
PreprintASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning
We propose ASGARD, a two-phase teacher-student pipeline for making RL-based UAV control resilient to action-space attacks.
PreprintReal-world useVideo-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models
We first train Qwen3-VL-8B with GRPO on a standard video dataset, and a second stage on Video-HopChain then raises the mean over eight video understanding and reasoning benchmarks from 55.4 to 57.9 and improves every one of them.
PreprintAD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control
We introduce AD-WM, an action-discriminative joint-embedding world model for counterfactual MPC.
PreprintMaskVLA: Visual Masking Against Trajectory Overfitting of Vision-Language-Action Model
To guide the model in effectively utilizing wrist camera information, we propose MaskVLA, a masking-based fine-tuning strategy.
PreprintReal-world useFIRM-WM: State-factorized factual-interventional recurrent modeling for reward-free visual planning
We introduce FIRM-WM (Factual--Interventional Recurrent World Model), a compact pixel world model designed around these two gaps.
PreprintMT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation
We present MT-WAM, which retains the original training objectives and adds complementary supervision for future two-dimensional point trajectories and visual features.
PreprintReal-world use