New Robotics & engineering papers
140 new Robotics & engineering papers appeared in the last 3 days. Pipette read every one, and these are the 60 with the broadest interest, a real step forward and careful claims. Each shows the sentence of its abstract that states the main result, exactly as its authors wrote it.
The best of the last 3 days
Dynamic flexible optical sensing based on liquid integrated circuits
Here, we propose a DFOS system based on liquid integrated circuits (LICs), integrating photosensitive ionic liquid (PIL) and stretchable liquid metal electrodes to achieve active curvature modulation for geometry-adaptive optical sensing.
Peer-reviewed journalReal-world useAn Analysis of Streaming Deep Reinforcement Learning for Adaptive Continual Learning in Robotics
In this paper, we present the first analysis of streaming deep reinforcement learning for adaptive continual learning in robotics.
PreprintClaims a big stepReal-world useAI-driven actuation-compatible tracking for closed-loop navigation of miniature robots in vivo
Here, we present an artificial intelligence (AI)-driven tracking system to overcome these challenges, enabling precise localization of millimeter-scale MWMRs in the presence of actuation fields.
Peer-reviewed journalReal-world useStabilizing telerobotic endobronchial imaging and interventions in breathing lungs with a helical brace
In this work, we present a reconfigurable helical brace that overcomes these challenges.
Peer-reviewed journalReal-world useFly, Drive, Reconfigure: A Modular Reconfigurable Aerial-Ground Platform for Field Operations
We present HARP, a Heterogeneous Aerial Robotic modules Platform in which independently deployable aerial robots physically reconfigure to compose their capabilities for field operations.
PreprintReal-world useFaBrick: An actuated fabric building platform for constructing scalable and versatile soft robots
Together, these results establish FaBrick as a scalable, versatile, and user-friendly foundation for building integrated soft robotic systems from a single material platform.
Peer-reviewed journalReal-world useGridSFM: A Foundation Model for Solving AC Optimal Power Flow
We introduce GridSFM, a framework that combines a pretrained foundation model across grid topologies with physics-informed fine-tuning for solving AC Optimal Power Flow (AC-OPF) at scale.
PreprintBold claims, read criticallyReal-world useRobots That Take Initiative: A Framework for Building and Evaluating Proactive Robots
We introduce a unified formalism for proactive robot assistance, organize it into three levels, and provide a framework to address the highest level of unprompted proactive assistance.
PreprintIntegrating Local Detail and Global Context: A Dual-Input Multi-Task Learning Framework for Bone Tumor Diagnosis
To address the limitations of existing single-view models, we present a dual-input, multi-task learning framework that, to our knowledge, is the first to apply bidirectional cross-modal attention between a lesion crop and the full radiograph for joint segmentation and subtype classification.
PreprintBold claims, read criticallyReal-world useAnthropomimetic Soft Robotic Forearm with Independently Articulated Carpal Bones Enabling Human-Like Adaptive Stiffness Modulability
These findings demonstrate that carpal bone morphology plays a dominant role in human wrist stiffness modulation and provide design principles for humanoid robot wrists.
PreprintReal-world useCode availableOctopus-inspired muscle activation unlocks high-DOF motion from a single actuator
Our results demonstrate a new actuation paradigm for achieving high DOF, offering a compact, efficient blueprint for soft robot design.
Peer-reviewed journalBold claims, read criticallyReal-world useRAPID: Robot Agentic Programming from Demonstrations
To leverage their potential for robot systems, this work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration.
PreprintTemperament Engineering: Designing Strategic Behavioural Diversity in Robot Swarms
This perspective proposes 'temperament engineering', a bio-inspired framework that treats the swarm's distribution of temperaments, rather than the individual controller, as the design object.
PreprintRACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning
We introduce RACaP, an agentic framework that moves coding to evolution and uses a Reasoning-and-Acting (ReAct) loop to call frozen, typed Policy APIs at deployment.
PreprintReal-world useHarnessPAI: An Evolving Harness for Physical AI
Across desktop robot arms, household robots, a robot vacuum, and a legged walking agent, HarnessPAI improves on both pure action models and code-as-policy baselines without retraining the underlying model: a 61.6-point gain over on LIBERO-PRO and a 27.2-point gain over WorldDreamer on RoboCasa atomic tasks.
PreprintBeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video
We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions.
PreprintReal-world useGPT-6-Astra Lights Up Embodied Navigation: Evaluation in Zero-Shot Vision-and-Language Navigation in Continuous Environments
First, \textbf{GPT-6-Astra achieves strong zero-shot navigation performance using only monocular RGB observations}.
PreprintBold claims, read criticallyHuman-in-the-Loop Geospatial Annotation for Rapid Dataset Construction in Field-Deployed UAV Systems
We present BirdsEye, which shifts expert annotation from images to the field: an operator records target locations in world coordinates using RTK positioning and calibrated projective geometry propagates each observation to all frames where the target is visible.
PreprintReal-world useRevolutionizing Diffusion MRI Microstructure Mapping via Global Inversion
We instead cast MM as a single global inverse problem, reconstructing the entire parameter field jointly from all measurements of a subject.
PreprintReal-world useWorld Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace.
PreprintUnlocking fast robotic locomotor propulsion through dynamic spine-leg synergy
By harnessing this synergy and its structural advantage, FLEXOR achieves a substantial increase in speed with a reduced cost of transport, outperforming state-of-the-art quadruped robots with flexible spines.
Peer-reviewed journalBold claims, read criticallyReal-world useOCC4M: Object-Centric 4D Memory for Spatiotemporal Reasoning in Long-Horizon Manipulation
We present OCC4M ("Occam"), an object-centric 4D memory that maintains persistent tracks in a shared world frame and explicitly represents temporal, motion, and containment relations.
PreprintTRACE: Interactive Bi-Directional Tracing of Monochrome Cables Amid Clutter
Evaluation with 110 physical experiments suggests that TRACE can increase the percentage of cable length correctly traced in complex scenarios (with up to 4 cables and 40 crossings) from ~60% with the strongest prior method, HANDLOOM 2.0, to ~90%, outperforming RT-DLO, Nano Banana Pro, and ChatGPT 5.2 as well.
PreprintReal-world useMorphIK: Morphology-Conditioned Neural Inverse Kinematics for Unknown Robots
We present MorphIK, a flow-matching model that solves inverse kinematics for revolute-joint-based kinematic chains it has never seen during training.
PreprintBold claims, read criticallyTraining-free Behavior Cloning
We introduce Behavior Predictive Control (BPC), which synthesizes policies without end-to-end policy training by combining an action-aware retrieval metric, a Hankel-based action-continuation prior, and a closed-form one-step residual correction.
PreprintReal-world useCoupled State-Space Modelling, Control, and Policy Distillation for Hybrid Rigid-Pneumatic Manipulators
The distilled policy settles 93-94 of goals with zero collisions, within a few points of its teacher, and runs inside the 5 ms control step where the MPC does not.
PreprintMorphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
We present Morphometric Imitation, a three-stage framework that transforms reconstructed HOIs into zero-shot sim-to-real visuomotor policies.
PreprintReal-world useOREN-X: Octree Residual Network for Real-Time Multi-Modal Mapping
We develop OREN-X, an online mapping method that uses an octree in 3D space as a shared data structure for indexing and storing a multi-modal field, capturing geometric, radiance, and vision-language information.
PreprintBold claims, read criticallyReal-world useFMCW-LIO: A Doppler LiDAR-Inertial Odometry
In the letter, we propose FMCW-LIO, a novel and robust LIO, leveraging intrinsic Doppler measurements from FMCW Doppler LiDARs.
Preprint with a published versionReal-world useNetwork Design against the Bullwhip Effect in Complex Supply Chains
We show that on a directed acyclic network every response from demand to orders is a sum over paths of products of nodal responses, so that the stability and stability margins of the network are determined by those of its individual nodes, and that a node which splits its orders across routes of different lengths gains margin it can spend on a higher gain and a faster response.
PreprintDAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models
We propose DAWN (Denoising and Alignment in World models for Noise-robustness), a noise-robust perception framework for legged locomotion, which builds noise robustness directly into a world model via two modifications: (1) feeding noisy depth to the encoder while keeping clean depth as the reconstruction target, forcing the model to implicitly denoise its input; and (2) applying contrastive learning to align the latent states of noisy and clean depth.
PreprintBold claims, read criticallyReal-world useLoRa Fluid Antenna Multiple Access
This paper advocates a new fluid antenna multiple access (FAMA) framework for LoRa, referred to as lora-FAMA, to provide spatial opportunities for LoRa multiuser communications.
PreprintReal-world useADM-Planner: LLM-Guided Long-Horizon Planning for Mobile Manipulators with Attention-Enhanced Dynamic Memory
To resolve this tension, we present an LLM-guided planning framework ADM-Planner with attention-enhanced dynamic memory (ADM).
PreprintCrossSafe: Towards Cross-Embodiment Latent Safety Filters
Our results show that a single policy, jointly trained across five manipulation tasks and four embodiments, exhibits zero-shot generalization to a held-out embodiment, reducing the nominal policy's collision rate.
PreprintReal-world useBioimpedance meets biomechanics: Wearable electrical impedance myography encodes fascicle and activation dynamics
These findings validate EIM as a robust, real-time biomechanical tool for sensing muscle kinematics and neuromuscular activity during functional movement, expanding its use in assistive robotic control and injury prevention.
Peer-reviewed journalBold claims, read criticallyReal-world useTactileStep: Sole Tactile Learning for Regulating Foot-Terrain Interaction in Humanoid Locomotion
Evaluated in simulation and on a Unitree G1 humanoid across diverse terrains, TactileStep reduces peak touchdown force by up to 48.8% and peak A-weighted impact noise by up to 30.1 dB over a strong perceptive baseline, while increasing stance contact area by up to 23.8%.
PreprintReal-world useWing-tail coordination balances efficient cruising and maneuvering flight in an avian-inspired flapping robot
A model-guided wing-tail coordination strategy enables NPU-Sparrow to balance cruise efficiency, longitudinal stability and maneuvering capability while expanding its flight-speed envelope and supporting agile maneuvers.
Peer-reviewed journalKoopman-Accelerated Model-Based Diffusion for Real-Time Robot Control
To address this limitation, this paper proposes bilinear Koopman model-based diffusion (BK-MBD).
PreprintReal-world useReVAMP: Vector-Accelerated Motion Planning for Kinematically-Constrained Systems via Reparameterization
We show that the planner can synthesize plans in microseconds to milliseconds for high dimensional systems (up to 20 dimensions), with complex constraints, up to 10x faster than the current state-of-the-art.
PreprintBold claims, read criticallyReal-world usePolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation
We present PolyUMI, an open-source platform for scalable visual--tactile--audio demonstration collection and robot deployment.
PreprintSafe Receding Horizon Mixed-Integer Differentiable Predictive Control for Degradation-Aware Battery Dispatch
We present a safe receding-horizon mixed-integer differentiable predictive control methodology for residential battery energy storage dispatch that combines neural-network speed with recursive feasibility guarantees.
PreprintReal-world useCoding Agents for Generalized Task and Motion Planning Problems
Overall, we find that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success (56% to 95% versus 47% for the planners, on the 16 environments where a planner is available).
PreprintContinuous Online Fault Detection for Mobile Robots via Adaptive Edge Models
These results validate the deployment of state-of-the-art anomaly detection on resource-constrained robotics through offline-to-online distillation.
PreprintReal-world useFree-Init: Scan-Free, Motion-Free, and Correspondence-Free Initialization for Doppler LiDAR-Inertial Systems
By fusing point-wise Doppler velocity with inertial measurements under non-inertial kinematics, the proposed framework, Free-Init, eliminates reliance on motion undistortion of LiDAR scans, excitation motions, and map correspondences during the initialization phase.
Preprint with a published versionBold claims, read criticallyReal-world useRepresentation World Model: Learning States, Transition and Executable Plans in Representation
We propose the Representation World Model (RWM), which learns states, transitions, and executable plans directly in representation space.
PreprintFrom Passive Execution to Active Exploration: Agentic Embodied Manipulation in Realistic Environments
To bridge this gap, we propose an agent-based active exploration framework that enables robots to dynamically interact with the environment rather than merely execute predefined instructions.
PreprintPlanning electric bus systems with solar photovoltaic integration using open transit data: A case study of the Dakar BRT
To address this gap, we present GTFS4EV, an open-source framework that uses publicly available General Transit Feed Specification (GTFS) data to simulate bus operations and evaluate electrification scenarios.
PreprintReal-world useAdaHVLA: Adaptive Harnesses for Long-Horizon Vision-Language-Action Execution
To bring these complementary capabilities together, we introduce AdaHVLA, an adaptive harness that refines code-based coordination policies through robot experience to better align agent reasoning and memory with VLA execution.
PreprintA Unified Frequency-Domain Model for Cascaded Filter-Interpolation Modulation in Tomographic Reconstruction
Here, we introduce a unified frequency-domain model that conceptualizes the combined effect of filtering and interpolation in the filtered backprojection (FBP) algorithm as a cascaded modulation process.
PreprintBold claims, read criticallyReal-world useFrom Target Selection to Digging: A Learning-Based Framework for Continuous Autonomous Excavation
We present a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers.
PreprintReal-world useReVNM: Learning-Based Visual Navigation from a Remote Camera
This paper presents the Remote Visual Navigation Model (ReVNM), which uses a single remote surveillance camera to serve as both an observation source and an implicit environmental map for visual navigation.
PreprintReal-world useOutcome-Sensitive Motion Search for Impact-Aware Dexterous Catching
Extensive simulation experiments demonstrate that our method effectively repairs task conditions where the teacher fails and enables the resulting IL policy to outperform the privileged RL teacher in both catching success and impact mitigation.
PreprintCALM: Current Aligned Link Manipulation for Single Arm Oversized Object Lifting
To address these challenges, we propose Current-Aligned Link Manipulation, a framework for learning long-horizon contact-rich manipulation using motor current as joint load related feedback.
PreprintDecoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs
We propose a framework that exposes backbone depth , action expert depth , and denoising steps as three jointly configurable compute axes in a VLA.
PreprintReal-world useEgoSpeedUp: Transferring Human Manipulation Tempo to Robot Policies
We introduce EgoSpeedUp, a framework that uses human manipulation as temporal supervision for robot imitation learning.
PreprintReal-world useHolistic co-design of fibre-optic distributed acoustic sensing and coherent communication
Here, we propose a holistically co-designed sensing and communication (CSAC) framework that incorporates multi-level optimization: (i) a dual-function pilot signal that shares time and spectrum resources; (ii) a simplified system using a single transmitter to simultaneously generate multi-tone pilot and subcarrier communication signals; (iii) a point-to-multipoint (P2MP) network that distributes pilot signals for multi-user frequency synchronization and enhanced acoustic sensing response.
Peer-reviewed journalBold claims, read criticallyReal-world useKnow Your Body: A Harness for Direct and Self-Improving Robot Control with VLMs
We introduce KnowBody, a harness that makes these action-relevant body relations explicit, queryable, and revisable while keeping the model weights frozen.
PreprintKeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization
We introduce KeyGen, a framework that learns canonicalized semantic 3D keypoints from point clouds and uses them as structured object-centric representations for policy learning.
PreprintReal-world useExact Factorisation and Fast Computation of Invertible Constant-Q Transforms
An exact factorisation combines spectral selection, conjugation, windowing, and reordering into a fixed map between one packed Fourier transform and the shorter band inverse transforms.
PreprintReal-world useActGaze: Learning Action-Grounded Gaze through Counterfactual Visual Interventions for High-Precision Manipulation
Extensive real-robot experiments on four high-precision robotic manipulation tasks demonstrate that ActGaze induces more focused visual attention on task-relevant regions and consistently outperforms the base VLA policy and other visual-grounding approaches.
PreprintReal-world use