New papers on Robots & autonomy
499 new papers on robots & autonomy in the last 7 days, within Robotics & engineering. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control
To this end, we propose {ForgetMimic}, the first motion-level unlearning method designed specifically for physical-world humanoid control.
PreprintClaims a big stepReal-world useCode availableCharacterizing Wildlife Response to Biomimetic and Conventional Underwater Vehicles
We present the first dataset comparing fish disturbance in response to a conventional thruster-driven AUV, a sea turtle inspired flipper-driven AUV, and a diver.
PreprintReal-world useGeneralizable Robotic Insertion with World Models
Our model-based approach trains a single world model on up to 90 insertion tasks with geometrically diverse parts, achieving 56% zero-shot success on unseen objects with unknown geometry compared to just 7% with a model-free baseline.
PreprintClaims a big stepReal-world useInsertAnything: Generalizable Contact-Rich Precision Insertion from Simulation to Reality
We present a reinforcement learning framework that trains insertion policies entirely in simulation for direct deployment without real-world demonstrations or policy fine-tuning.
PreprintClaims a big stepReal-world useAthenaZero: A low-inertia, bimanual robot for dynamic manipulation
AthenaZero is a bimanual manipulator designed to minimize inertia without compromising control authority.
Preprint with a published versionDr-LiSA: Direct Radar-Lidar Scan Alignment for Localization
This paper introduces Dr-LiSA, a first-of-its-kind direct method for localizing 2D spinning radar intensity measurements in against 3D lidar maps.
PreprintClaims a big stepReal-world useRiverVLN: Phase-Grounded Temporal Vision--Language Navigation for Unmanned Surface Vehicles
We introduce RiverVLN, to our knowledge the first benchmark designed for long-horizon USV VLN under continuous riverine motion, and PGT-NAV, a phase-grounded temporal navigation framework for USVs.
PreprintClaims a big stepReal-world useAn Analysis of Streaming Deep Reinforcement Learning for Adaptive Continual Learning in Robotics
In this paper, we present the first analysis of streaming deep reinforcement learning for adaptive continual learning in robotics.
PreprintClaims a big stepReal-world useRegenHarness: A Robot Agent Harness with Evidence-Gated Recursive Self-Improvement
To our knowledge, we are the first to introduce an evidence-gated recursive self-improvement (RSI) protocol for embodied robotic agents.
PreprintBold claims, read criticallyClaims a big stepReal-world useFrom Instrument-Mounted Demonstrations to In-Vivo Execution: Learning Bimanual Laparoscopic Appendectomy Without Robot-Collected Demonstrations
The results show that demonstrations recorded from a surgeon's own instruments are sufficient to train, select and deploy a bimanual surgical policy in vivo.
PreprintReal-world useREBOOT: From Failure to Recovery - A Dataset and Benchmark for Precision Assembly
We introduce REBOOT (Recovery Episode Benchmark for Off-nominal Trajectories), the first robot manipulation benchmark designed around failure as a first-class signal.
PreprintAI-driven actuation-compatible tracking for closed-loop navigation of miniature robots in vivo
Here, we present an artificial intelligence (AI)-driven tracking system to overcome these challenges, enabling precise localization of millimeter-scale MWMRs in the presence of actuation fields.
Peer-reviewed journalReal-world useStabilizing telerobotic endobronchial imaging and interventions in breathing lungs with a helical brace
In this work, we present a reconfigurable helical brace that overcomes these challenges.
Peer-reviewed journalReal-world useHumynexSurg-1: A Curated Expert Liposuction Dataset
HumynexSurg-1 is the first release: a master liposuction surgeon performing on porcine abdominal tissue while narrating every decision, recorded with synchronized suction pressure, six-axis hand force/torque, top-down RGB-D video, side video and a lavalier microphone -- 14 episodes, 42,738 frames, 35.6 minutes, 356 utterances of which 95% compile into a liposuction-specific label schema.
PreprintReal-world useFly, Drive, Reconfigure: A Modular Reconfigurable Aerial-Ground Platform for Field Operations
We present HARP, a Heterogeneous Aerial Robotic modules Platform in which independently deployable aerial robots physically reconfigure to compose their capabilities for field operations.
PreprintReal-world useInternW0: A Foundational Physical World Model for Efficient Real-World Interactions
We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi-frequency processing, and local physical modeling under partial observations and external influences.
PreprintBold claims, read criticallyClaims a big stepReal-world useA foldable small-scale soft electromagnetic robot for multimodal navigation in confined and unstructured environments
Here, we introduce a small-scale, compact, foldable, and robust soft electromagnetic robot with more than nine locomotion modes designed for such a scenario.
Peer-reviewed journalReal-world useEstimation and Control of Tensegrity Manipulator Kinematics based on Strut Inclination Angles
To the best of our knowledge, this work presents the first experimental demonstration of a real-time IMU-based shape estimation method on a full-scale tensegrity manipulator and demonstrates posture control using a simple Proportional-Integral (PI) controller.
PreprintDEXTERA: From a Single Image to Deployable Dexterous Manipulation via Real-to-Sim-to-Real
We present DEXTERA, an automated real-to-sim-to-real framework that transforms a single RGB image into deployable policies for dexterous manipulation across four unified stages: (1) single-image scene factorization into a static Gaussian background and interactive rigid or articulated assets with VLM-inferred physical parameters; (2) metric scene global alignment, object canonicalization, and morphology-balanced robot calibration; (3) scalable simulator task primitive construction, VR teleoperation, and object-centric trajectory synthesis; and (4) a shared multimodal policy interface supporting both imitation learning and reinforcement learning.
PreprintReal-world useFaBrick: An actuated fabric building platform for constructing scalable and versatile soft robots
Together, these results establish FaBrick as a scalable, versatile, and user-friendly foundation for building integrated soft robotic systems from a single material platform.
Peer-reviewed journalReal-world useMotionForge: A Data Generation Pipeline and Large-Scale Benchmark for Long-Horizon Manipulation of Dynamic Objects with Domain Shifts
To bridge these gaps, we introduce MotionForge, the first large- scale simulation benchmark and data-generation pipeline tailored to jointly evaluate domain shifts and long-horizon interaction in dynamic manipulation.
PreprintTip Manipulation in Soft Everting Robots via Wall Retraction and Deployable Fingers
We introduce a tip-manipulation and multi-tool deployment strategy for soft everting robots based on wall retraction, implemented with a base roller assembly that independently meters membrane flow in the outer wall while a tail spool regulates growth in the internal tail section.
PreprintReal-world usePhrase-Level Robotic Guqin Performance: Bimanual Motion Planning and Audio-Tactile Interaction Monitoring
In this work, we present a physical heterogeneous dual-arm robotic system for phrase-level autonomous guqin performance.
PreprintDexterous Robot Manipulation from Human Demonstrations via Contact-Anchored Retargeting and Residual Policy Learning
The pipeline reconstructs 25,454 single-hand trajectories (success 7.3% -> 59.3%) and 25 dual-hand tasks (16.0% -> 62.4%) with one shared policy per setting, transfers one human dataset to four morphologically distinct robot hands (+62.4 pp), and executes four contact-rich bimanual tasks on physical hardware with zero real-robot training data.
PreprintReal-world useCode availableRobots That Take Initiative: A Framework for Building and Evaluating Proactive Robots
We introduce a unified formalism for proactive robot assistance, organize it into three levels, and provide a framework to address the highest level of unprompted proactive assistance.
PreprintBend the Clock: Predicting Ahead to Beat Latency in Event-Based Object Detection
To address this, we introduce ChronoFuse, a causal availability-time detector that predicts object states for when its output becomes available rather than for when its input was observed.
PreprintReal-world useEgoWild2Dex: Learning Dexterous Robotic Manipulation from In-the-Wild Human Experience
To address these issues, we introduce EgoWild2Dex, which transfers in-the-wild ego-human experience to dual-arm robots with dexterous hands by jointly aligning unstable egocentric views and human motions with robot observations and actions, respectively.
PreprintReal-world useAnthropomimetic Soft Robotic Forearm with Independently Articulated Carpal Bones Enabling Human-Like Adaptive Stiffness Modulability
These findings demonstrate that carpal bone morphology plays a dominant role in human wrist stiffness modulation and provide design principles for humanoid robot wrists.
PreprintReal-world useCode availableThe Cartesian Hand: In-Hand Manipulation with All-Linear Fingers
We introduce the Cartesian Hand, a 7-DoF end-effector that rethinks dexterous manipulation by combining independent grasping and relative manipulation within a single end-effector using only linear motion.
PreprintReal-world useStructured World-State Reasoning for Agentic Robotic Search
WORLDS achieves 51.8% navigation success across all 5,311 CityNav test episodes, the highest reported success rate, exceeding the previous published best by 15.7 percentage points under an OSM-only, high-resolution orthographic protocol.
PreprintEmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics
Together, our framework uses coding agents to solve complex robotics tasks and turn verified solutions into scalable supervision for robot policies.
PreprintReal-world useDexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation
We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling.
PreprintDiverse and Adaptable Arm Coordination for Octopus-Crawling via Diffusion-Based Uncertainty-Aware Optimization
The main contributions include: (i) a symmetry-structured policy representation that folds radially equivalent controllers into a canonical directional sector, (ii) an online black-box optimization strategy, the DUO algorithm, that discovers and retains diverse coordination modes, and (iii) a control editing technique that adapts existing controllers to novel actuator constraints without retraining.
PreprintBold claims, read criticallyClaims a big stepDynamic, Decentralized Spatial Code Reuse for OCDMA LiDAR in Robot Swarms
We propose a decentralized protocol in which robots dynamically reassign spatial reuse codes based on a live, beacon-maintained interference-neighborhood graph, and prove that the number of codes required grows as under constant robot density -- an unbounded improvement over the growth of static assignment.
PreprintReal-world useOctopus-inspired muscle activation unlocks high-DOF motion from a single actuator
Our results demonstrate a new actuation paradigm for achieving high DOF, offering a compact, efficient blueprint for soft robot design.
Peer-reviewed journalBold claims, read criticallyReal-world useRAPID: Robot Agentic Programming from Demonstrations
To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration.
PreprintMATE: Multi-Agent Virtual Teleoperation Platform for Humanoid Collaboration Data Collection
In this work, we introduce MATE, a Multi-Agent virtual TEleoperation platform for humanoid collaboration data collection that enables multiple geographically distributed operators to simultaneously control whole-body humanoids in a shared physics-based environment.
PreprintTemperament Engineering: Designing Strategic Behavioural Diversity in Robot Swarms
This perspective proposes 'temperament engineering', a bio-inspired framework that treats the swarm's distribution of temperaments, rather than the individual controller, as the design object.
PreprintRACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning
We introduce RACaP, an agentic framework that moves coding to evolution and uses a Reasoning-and-Acting (ReAct) loop to call frozen, typed Policy APIs at deployment.
PreprintReal-world useOmnidirectional Amphibious Locomotion via Internal Mass Actuation
We present MARBLE, a fully enclosed omnidirectional amphibious rolling robot driven entirely by internal mass redistribution.
PreprintReal-world useDemonstration Synthesis from a Single Scan via Gaussian Splatting for Visuomotor Policy Learning
This paper introduces GaussianFactory, a high-fidelity data engine that mass-produces demonstrations with a single video scan as its only human input and no physics engine in the generation loop.
PreprintReal-world useA High-Payload Wall-Climbing Robot Using Passive Bistable Suction Cups
This work presents a novel high-payload wall-climbing robot that utilizes passive bistable suction cups to generate adhesion without needing to be pushed into the wall.
PreprintReal-world useHIGenNTO: Scalable Humanoid Interaction Generation via Noise-Space Trajectory Optimization
We present HIGenNTO, a framework that synthesizes humanoid-scene interaction motion references by optimizing the initial noise of a pretrained text-conditioned motion model under sparse spatiotemporal and scene constraints.
PreprintReal-world useARSTAG: An Agentic Real2Sim2Real System for Task-Specific Robot Data Generation
We present ARSTAG, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data.
PreprintReal-world useVisuomotor Robotic Pruning in Planar Orchards Using Hybrid Reinforcement Learning
We introduce an end-to-end pipeline to learn a closed-loop visuomotor controller for robotic pruning.
PreprintReal-world useControlling Collectives of AI Agents in Reasoning Space with Spatial Transformers
We propose COMPASS, a scalable, decentralized multi-robot architecture for controlling large collectives of agentic robots with reasoning space feedback control.
PreprintBold claims, read criticallyReal-world useHarnessPAI: An Evolving Harness for Physical AI
Across desktop robot arms, household robots, a robot vacuum, and a legged walking agent, HarnessPAI improves on both pure action models and code-as-policy baselines without retraining the underlying model: a 61.6-point gain over on LIBERO-PRO and a 27.2-point gain over WorldDreamer on RoboCasa atomic tasks.
PreprintBeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video
We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions.
PreprintReal-world useGPT-6-Astra Lights Up Embodied Navigation: Evaluation in Zero-Shot Vision-and-Language Navigation in Continuous Environments
First, \textbf{GPT-6-Astra achieves strong zero-shot navigation performance using only monocular RGB observations}.
PreprintBold claims, read criticallySynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations
SynthDemo-RL, with 50 synthesized trajectories per task and no new human demonstrations, rescues all 27 and reaches average success rates of 97.8% and 97.1% on the Position and Task axes of LIBERO-PRO, respectively.
Preprint