pipette
ESEspañol

TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation

Hengyan Liu, Wenlve Zhou, Bo Yue, Yongyi Su, Ruixiang Wang, Zhanqi Zhang, Dekun Lu, Wei Gao, Xiaofen Xing, Kui Jia

Preprint

In the authors' words

Reactive vision--language--action (VLA) models struggle with long-horizon manipulation when visually similar observations can correspond to different actions depending on the task stage or interaction history. We refer to this ambiguity as task-state aliasing and introduce TaskAnchor, a lightweight adapter that grounds pretrained VLAs in execution history. TaskAnchor combines history-conditioned visual refinement with a milestone-supervised task-state coordinate, a scalar representing the semantic stage of execution. These signals are injected through the native visual and language interfaces, respectively, without introducing an explicit planner or modifying the action-generation mechanism. On RMBench, TaskAnchor achieves approximately 4.9--5.5 the average success rates of the published and X-VLA baselines, with consistent gains on RoboMemArena and real robots. The added latency is only 2.08 ms per action chunk for .

Main resultThe abstract does not state a limitation.

Appeared: Tuesday, September 22. arXiv. Preprint, not yet peer-reviewed.