AI agents do more of the work in model development, but humans still make the decisions
Atria's development study of its 744B-parameter Atria Dawn Preview found agents executed more steps while humans retained over 85% of final decisions.
The Atria team built Atria Dawn Preview, a 744-billion-parameter mixture-of-experts agentic model trained with real execution environments, leading 5 of 16 benchmarks including AutomationBench, CyberGym, and MLE-Bench Lite while trailing on GDPval and SWE-Bench Pro. Across the project, AI was used in 96.5 percent of tasks, the median ratio of agent actions to human inputs rose from 11 to 28.5 over four weeks, and 151 of 455 completed AI-assisted tasks were rated infeasible without AI. Humans still made 85.5 percent of method and parameter decisions versus AI's 9.2 percent, and chose goals and scope in 93.4 percent of cases, even among AI-dependent tasks (95.4 percent). The team warns of a rubber-stamp oversight risk as agent work chains exceed human review capacity, situating findings in the recursive self-improvement debate involving Anthropic, OpenAI's GPT-5.6 Sol, and Google DeepMind's Dream-RSI.