TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
Researchers introduce TANGO, a whole-body vision-language-action model enabling zero-shot language-guided humanoid navigation on the Unitree G1 robot.
TANGO addresses humanoid navigation in cluttered indoor environments by predicting 29-DoF joint-space actions directly from natural-language instructions and egocentric RGB observations, going beyond 2D path planning. It is described as the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in clutter. The model is trained entirely in simulation via a pipeline combining global path planning, kinematic whole-body motion generation, obstacle-aware motion editing, and RL-based tracking. In simulation it achieves state-of-the-art vision-language navigation performance, and it transfers zero-shot to a Unitree G1 humanoid without any real-world navigation data or training data. The two source reports (Hugging Face daily papers, 2026-09-07, and arXiv cs.AI/cs.LG/cs.CL, 2026-09-08) are fully consistent; no disagreements or additional figures were reported.
- TANGO is a whole-body vision-language-action (VLA) model for humanoid navigation in cluttered indoor environments.
- It predicts 29-DoF joint-space actions directly from natural-language instructions and egocentric RGB, rather than 2D path planning.
- It is described as the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in clutter.
- It is trained entirely in simulation using a pipeline combining global path planning, kinematic whole-body motion generation, obstacle-aware motion editing, and RL-based tracking.
- In simulation it achieves state-of-the-art vision-language navigation performance.
- It deploys zero-shot on a Unitree G1 humanoid without any real-world navigation data or training data.
- Both source reports (dated 2026-09-07 and 2026-09-08) describe the work identically, with no conflicting figures.
Coverage timelineoldest first · each row is one article
- · 8d agoTANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
Hugging Face daily papers· 27
Researchers introduce TANGO, a whole-body vision-language-action model enabling zero-shot language-guided humanoid navigation on the Unitree G1 robot.