GPT-6-Astra Lights Up Embodied Navigation: Evaluation in Zero-Shot Vision-and-Language Navigation in Continuous Environments
In the authors' words
We investigate whether GPT-6-Astra, a general-purpose foundation model, can navigate unfamiliar environments using its own perception, reasoning, and decision-making capabilities. Our evaluation focues on zero-shot vision-and-language navigation in continuous environments (VLN-CE) through a minimal interface in the Codex harness, aiming to unleash GPT-6-Astra's full potential for navigation. Using monocular RGB, GPT-6-Astra decides when to observe, how to move, and when to stop, without navigation-specific fine-tuning, a trained waypoint predictor, or a pre-built scene map. Our evaluation yields four main findings. First, \textbf{GPT-6-Astra achieves strong zero-shot navigation performance using only monocular RGB observations}. On the common-adopted zero-shot R2R-CE benchmark, ultra reasoning achieves a success rate of \textbf{79.0%}, exceeding the strongest reported zero-shot and supervised success rates by \textbf{13.0} and \textbf{6.9} percentage points, respectively. Second, \textbf{GPT-6-Astra advances multi-stage language instructions into coherent, adaptive navigation} by grounding spatial relations, tracking task progress, and revising its actions. Third, \textbf{reliable route execution and goal verification remain challenging, even with ultra reasoning}. Plausible local landmark matches do not consistently lead to correct task completion. Fourth, \textbf{these capabilities motivate rethinking the role of embodied learning}. Future VLN research should build on foundation models to advance generalizable and reliable embodied intelligence.
Appeared: Friday, September 25. arXiv. Preprint, not yet peer-reviewed.
Authors' comment: Technical report