NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents
NeutronGym grades LLM agents on neutron instrument design via McStas ray-tracing, with RL lifting Qwen3-8B from 11% to 77%.
NeutronGym is presented as the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces the result, and a level-resolved ladder grades syntax, runtime, structure, and science without any LLM judge. Seven models reproduce at most 7 of 16 McStasBench tasks from published instruments, none retrieves a reference, and none meets an improvement target. Reinforcement learning on the environment's reward takes Qwen3-8B from 11% to 77% of held-out instances, past an untrained Qwen3-32B, while frontier models solve 98-99%. The gain collapses by 60 points without the ladder's partial credit.
- First executable benchmark for neutron instrument design, no LLM judge
- Frontier models solve 98-99% of the task domain
- RL training lifts Qwen3-8B from 11% to 77% on held-out instances
- Improvement collapses by 60 points without partial-credit grading ladder
Full article246 words · extracted from arxiv.org · click to collapse
Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with. We introduce NeutronGym, to our knowledge the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge. Procedural families supply unlimited instances of a fixed layout whose design parameters the agent must set, with held-out parameter regimes; a curated slice, McStasBench, adds 16 tasks from published instruments behind memorization probes and a sandbox. Seven models reproduce at most 7 of the 16, none retrieves a reference, and none meets an improvement target. The environment also trains. Reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B, and the recipe holds, at one seed each, on three further gated families. The analysis says what that gain is. Without the ladder's partial credit it collapses by 60 points. From reward alone the trained model reaches what a classical optimizer reaches, at the agent's simulation budget, only when handed the closed-form physics (77% against 81%, a gap that does not separate at this size), while frontier models still solve 98-99%. Getting a trustworthy result meant failing four task designs that no-model baselines could solve, and we release the probes that found them.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.03631