New papers on Agents & reasoning
167 new papers on agents & reasoning in the last 7 days, within AI & machine learning. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
Recursive self-improvement of AI research agents
Together, these results show that an AI research agent can improve its own research efficiency through recursive self-improvement, and that these gains transfer to tasks and domains the loop never encountered.
PreprintClaims a big stepMedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution
We introduce MedRSI, the first recursive self-improvement framework for medicine, which continuously transforms diagnostic failures into new clinical capabilities through tool composition and task-specific model training.
PreprintBold claims, read criticallyClaims a big stepReal-world useCode availablePUBG Ally: A Conversational Embodied Agent as an AI Teammate
We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate.
PreprintReal-world useBI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
BI-Agent achieves substantial accuracy gains of up to 40 percentage points with vanilla LLMs, and post-trained BI-Agent yields gains of up to 30 points.
PreprintReal-world useCode availableGameDirector: Decoupling Gameplay Logic from Rendering for Player-Configurable Game World Models
To bridge these paradigms, we introduce GameDirector, the first agentic framework that decouples rule-based gameplay logic from visual rendering.
PreprintAgents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control
Across 1,750 cells, 1,419 (81.1%) satisfy the verifier, and 808 (46.2%) also survive every stricter filter: visible, localized, typeface-matched, original value gone document-wide.
PreprintReal-world useLabFactory: Building and Evaluating Executable AI Labs
We present LabFactory, a framework in which an AI builder turns a scientific brief into an executable AI lab: a task-specific solver that integrates models, knowledge resources, tools, and a controller behind a fixed interface.
PreprintReal-world useProvably Complete Generalized Planning with LLMs
Here, we present an approach for automatically generating generalized plans in Lean together with proofs of their completeness relative to a specification of the domain constraints provided as input.
PreprintBold claims, read criticallyClaims a big stepCAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
These results establish incentive robustness as a distinct challenge for delegated agents, diagnose how it fails, and show that targeted interventions can substantially improve it.
PreprintReal-world useSelf-Organizing Agent Teams Learn to Reason Together
More broadly, these results suggest that organization itself can become an agent capability: agent teams can learn how to reason together and produce solutions their members could not reach independently.
PreprintBold claims, read criticallyPropose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors
Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge.
PreprintAPEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction
We present APEXA, a deployed multi-agent framework (61 tools over heterogeneous compute, run as a single reasoning loop) automating calibration and integration from natural language at a major light source.
PreprintReal-world useCode availableAstraLOD3: Zero-shot multimodal agentic reconstruction of LOD3 building models
The results demonstrate that structured LOD3 reconstruction can be formulated as a constrained agentic process rather than as a fixed pipeline.
PreprintSocioVerse2: A Longitudinal Dynamic Social Simulation Framework under a Human-AI Co-evolutionary Paradigm
We present SocioVerse2, which extends SocioVerse 1.0 into a human-AI co-evolutionary paradigm built from two loops and one infrastructure.
PreprintTrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization
Here, we introduce TrialAtlas, a memory-augmented multi-agent research organization for CDP that mirrors this collaborative process by coordinating specialized agents for literature synthesis, competitive trial intelligence, regulatory precedent analysis, and integrated reasoning over trial design and development risk.
PreprintReal-world useHarness-Zero: Harness Distillation via Agent-as-Harness
We introduce Harness-Zero, which enables harness distillation through agent-as-harness.
PreprintThe Last Human Gate: Forward Deployed Engineering for Governance Automation
These results support the technical feasibility of replacing human execution of specified governance-review tasks with agents and software.
PreprintReal-world useCode availableEmergent Intelligence: Resonant Oscillators Produce Proactive Adaptive Behavior
With no training, supervision, parameter tuning, or controller, the composite switches on its own between exploratory spiral search and exploitative tracking, finding both first-degree symmetry and second-degree groups.
PreprintBold claims, read criticallyCode availableDesigner-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience.
PreprintReal-world useScreen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale
We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment.
PreprintReal-world useWhere Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents
Across 25,930 episodes spanning nine recent models, three production agent harnesses, two contract variants and fifteen recovery conditions, the answer depends on the fault.
PreprintReal-world useLive Assistant: Learning Whether, When, and Whom to Assist in Real-World Live Social Streams
We introduce \liveassistant, a framework for mixed-initiative, role-conditioned assistance that formulates livestream interaction as four coupled decisions: whether to act, when to act, whom to address, and what to communicate.
PreprintReal-world useImpact Is Not Invalidation: Ask About the Claim, Not the Diff
Asked instead whether one specific claim still holds, the same models on the same diffs reach 0.705 to 0.974.
PreprintKITE: Scaling Jev Population Experiments with Sparse Flagship Calibration
This architecture offers a route to screening candidate interventions before human trials, multi-country content audits, and uncertainty-aware policy comparison at the cost of a few thousand kernel calls with sparse flagship anchors.
PreprintCode availableMind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation
Most tellingly, an agent predicts what its partner believes about it 72.5% of the time while that partner's belief is itself correct only 51.2% of the time: the agents track the conversation far better than they track the mind behind it.
PreprintVerifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
VHD-Play reverses this dependency by sampling and solving a mathematical model before a corpus-grounded setter renders its decision process as stateful tools.
PreprintDrGait: Biomechanically Grounded Visual Reasoning for Interpretable Clinical Gait Analysis
By anchoring VLM's reasoning in verifiable geometric and temporal measurements, DrGait reduces hallucinations, achieving competitive diagnostic accuracy while generating transparent and audit-ready clinical reports.
PreprintReal-world useAgentic Detection of Online Conspiracies
Evaluating our framework on a manually-annotated adversarial dataset, we find that context-aware workflows consistently outperform text-only classification and that the agentic framework performs significantly better than other frameworks and settings, including a non-agentic model exposed to the same contexts available to the agent.
PreprintQwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents
Qwen-Planner-Agent achieves the best overall performance among all evaluated models and systems on MobilePA-Bench, improving over its base model across tool use, memory, skills, and sub-agent coordination.
PreprintReal-world useScaling Discovery through Test-Time Communication
We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward.
PreprintGoverned AI-Agent Coordination for Dementia Care: Architecture, Safety Contracts, and Evidence-Derived Workflow Verification
GCAC satisfies all 18 contract oracles with zero policy-violating tool calls and correctly preserves obligations, rejects stale state, creates human hand-offs, and records workflow closure.
PreprintReal-world useAgent-Editing World Model: Rethinking World Modeling for LLM Agents
We propose the Agent-Editing World Model (AEWM), which models how reasoning and actions shape future task progress rather than simulating tool responses.
PreprintJust-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Across ALFWorld, WebShop, and -bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively.
PreprintBeyond Spatial Benchmarks: From Spatial Reasoning to Navigation
Our analysis reveals a gap between benchmark-oriented spatial specialization and navigation performance, and shows how aligning spatial supervision with navigation goals, phases, and decision learning improves navigation.
PreprintCode availableEmergi-PersonaOS: A Persona Agent Operating System for Situational Adaptation and Controllable Evolution
This work provides a computational framework for persona agents to maintain individual continuity, produce situation-specific expression, and develop through experience over sustained interaction.
PreprintBold claims, read criticallyFrom Certain Doom to Survival: Agent-Driven Self-Governance in LLM Agent Societies
Our results show that executable governance improves the space of possible interventions for agents, but survival depends on whether agents discover the right institutional mechanisms in time.
PreprintWorld State Generator
Across 7 public benchmarks, WSG raises end-to-end success for two open models near 30B parameters over prompting and brings to the level of proprietary model.
PreprintAutomatic multimodal UX improvement recommendations from LLM agent user simulations
We present AMUSER, a multimodal framework which simulates user behaviour and automatically generates prioritised UX improvement recommendations from resulting data.
PreprintReal-world useTowards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction
To overcome these barriers, we propose the MaP (stands for "Masked Trajectory Prediction"), a unified framework that seamlessly harmonizes divergent GUI navigation tasks.
PreprintWho Holds the Pen? Let Specifications, Not Agents, Sign Off
Experiments on guideline-following and artifact-generation tasks show that specifications can serve not merely as behavioral guidance, but as authority over compliant execution and completion.
PreprintMintAct: A Unified Visual Agent for Digital Environments
We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales.
PreprintReal-world useSAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance
Motivated by this analysis, we propose SAGE (Structural Admissibility-Guided Exploration), a unified framework that injects structural guidance to alleviate exploration bias and compounding bias in long-horizon reasoning.
PreprintCode availableA Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning
We introduce a generic, fully differentiable neuro-soft-symbolic framework that connects visual perception and task planning within a single computational graph.
PreprintQwen3.8-Omni: Towards Native Omni-Modal Agents
We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity.
PreprintBold claims, read criticallyReal-world useSocial Influence and the Allocation of Scientific Attention in AI Populations
The results show how a simple information rule shapes the volume, breadth and distribution of scientific attention in an artificial population.
PreprintCoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop
We present CoLearn, an interactive, agentic tutor that supports an iterative tutoring loop: the learner practises, and the system builds an evidence-grounded memory of the learner's mastery and misconceptions.
PreprintReal-world useVerify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale
We present EvoPilot, a human-gated method for long-horizon online autoresearch.
PreprintReal-world useTimeEvo: Failure-Driven Self-Evolution of a Time Series Agent
Experiments on ten time series QA tasks and three backbones show that TimeEvo, starting from an empty library, improves accuracy on every task and every backbone, and that a library grown on a cheap model still gains when it is installed into stronger ones.
PreprintCode availableBaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data Pipelines
This paper introduces BaseCamp, a novel agentic AI framework for automating the decision layer of DNA sequencing pipelines.
PreprintBold claims, read criticallyReal-world usePlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking
We present PlaceReasoner-Beta, a verifier-guided multi-agent framework that reformulates macro placement as a closed-loop reasoning problem rather than black-box optimization.
PreprintReal-world use