ZeroHour
Story · 1 source · 1 articlefirst updated ()1

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

infoAI researchimportance 28
What's new: Initial merged summary (no previous story). The arXiv report (2026-09-08) restates the Hugging Face daily papers report (2026-09-07) with no new results, figures, or corrections; the only differences are minor wording variants (e.g., 'without any real-world navigation data' vs. 'without any real-world navigation training data'), which are treated as equivalent.
Merged summary · glm-5.3-flash · rewritten as coverage arrives

Researchers introduce TANGO, a whole-body vision-language-action model enabling zero-shot language-guided humanoid navigation on the Unitree G1 robot.

TANGO addresses humanoid navigation in cluttered indoor environments by predicting 29-DoF joint-space actions directly from natural-language instructions and egocentric RGB observations, going beyond 2D path planning. It is described as the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in clutter. The model is trained entirely in simulation via a pipeline combining global path planning, kinematic whole-body motion generation, obstacle-aware motion editing, and RL-based tracking. In simulation it achieves state-of-the-art vision-language navigation performance, and it transfers zero-shot to a Unitree G1 humanoid without any real-world navigation data or training data. The two source reports (Hugging Face daily papers, 2026-09-07, and arXiv cs.AI/cs.LG/cs.CL, 2026-09-08) are fully consistent; no disagreements or additional figures were reported.

  • TANGO is a whole-body vision-language-action (VLA) model for humanoid navigation in cluttered indoor environments.
  • It predicts 29-DoF joint-space actions directly from natural-language instructions and egocentric RGB, rather than 2D path planning.
  • It is described as the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in clutter.
  • It is trained entirely in simulation using a pipeline combining global path planning, kinematic whole-body motion generation, obstacle-aware motion editing, and RL-based tracking.
  • In simulation it achieves state-of-the-art vision-language navigation performance.
  • It deploys zero-shot on a Unitree G1 humanoid without any real-world navigation data or training data.
  • Both source reports (dated 2026-09-07 and 2026-09-08) describe the work identically, with no conflicting figures.
VendorsUnitree
OrganizationsUnitree
AI modelsTANGO

Coverage timeline

  1. · 8d ago
    Hugging Face daily papers· 27
    TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

    Researchers introduce TANGO, a whole-body vision-language-action model enabling zero-shot language-guided humanoid navigation on the Unitree G1 robot.