ZeroHour
Story · 2 sources · 2 articlesfirst updated ()

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

infoAI researchimportance 42
What's new: Initial merged story created from two reports: Hugging Face daily papers (2026-09-10) and arXiv listing (2026-09-11). Both sources report identical claims and figures with no discrepancies, so the summary reflects the agreed facts.
Merged summary · glm-5.3 · rewritten as coverage arrives

A single omnimodal masked-diffusion VLA model unifies action, dynamics, and goal prediction, reaching 78.4% average success on Franka Research 3 manipulation tasks and up to 29.2x faster action decoding.

Dynin-Robotics is built on the Dynin-Omni omnimodal masked-diffusion backbone and represents language, observations, goals, and actions as discrete tokens within one shared trajectory model. A single model learns four objectives: action prediction, action-conditioned next-observation (dynamics) prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction. The shared trajectory interface enables test-time scaling via goal prediction, action-candidate evaluation, and joint action/future-state refinement. The model is continually pretrained on approximately 1.33 million trajectories from 48 Open X-Embodiment datasets. It achieves competitive results on LIBERO and zero-shot LIBERO-Plus, a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot, and up to 29.2x faster model-side action decoding from a block-parallel implementation. Both the Hugging Face daily papers listing and the arXiv (cs.AI/cs.LG/cs.CL) report state identical figures; no source disagreement.

  • Single omnimodal masked-diffusion model (Dynin-Omni backbone) learns action prediction, action-conditioned next-observation prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction.
  • Continually pretrained on ~1.33 million trajectories from 48 Open X-Embodiment datasets.
  • 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot.
  • Competitive performance on LIBERO and zero-shot LIBERO-Plus benchmarks.
  • Block-parallel implementation accelerates model-side action decoding by up to 29.2x.
  • Enables test-time scaling through goal prediction, action-candidate evaluation, and joint action/future-state refinement.

Coverage timeline

  1. · 5d ago
    Hugging Face daily papers· 42
    Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

    Dynin-Robotics unifies action, goal, and dynamics prediction in one omnimodal masked-diffusion VLA model, reaching 78.4% success on Franka Research 3 manipulation tasks.

  2. · 4d ago
    arXiv cs.AI / cs.LG / cs.CL· 35
    Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

    Dynin-Robotics omnimodal diffusion vision-language-action model unifies action, dynamics, and goal prediction, reaching 78.4% success on Franka tasks.