Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents
Omni-Decision plans omni-modal agent tasks from a compact evidence ledger instead of noisy history.
Omni-Decision is an omni-modal agent that replaces accumulating conversation history with an evidence ledger of missing, confirmed, and conflicting facts. A critic filters noisy video, audio, web, and computation observations, and the planner is further trained with supervised fine-tuning and decision-level reinforcement learning. It reaches 81.4% accuracy on OmniGAIA at roughly 43% of Gemini-3.1-Pro's cost per question and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model. Swapping the planner causes a larger drop than swapping the perception backend.
- Evidence ledger replaces growing noisy multimodal dialogue history
- A critic passes only usable observations to the planner
- 81.4% on OmniGAIA at about 43% of Gemini-3.1-Pro cost
- 65.0% on WorldSense, matching the strongest end-to-end model
Full article178 words · extracted from huggingface.co · click to collapse
Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict. A critic reads each noisy observation and passes only the usable content to the ledger, discarding the rest, so the planner works from a compact context throughout the task. Each run records the state, action, and verdict at every step, and supervised fine-tuning and decision-level reinforcement learning on these trajectories further improve the planner. Omni-Decision achieves state-of-the-art accuracy of 81.4% on OmniGAIA at approximately 43% of Gemini-3.1-Pro's cost per question, and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2607.11433