Reflection’s Beam trails top open models on coding tests but claims lower inference compute
Reflection AI’s Beam, a 501B open-weight coding model, trails top open models but claims lower inference compute.
Reflection AI announced Beam, a 501-billion-parameter open-weight mixture-of-experts model for coding and agents, with 23 billion parameters active per token. It plans to release the weights under Apache 2.0 later this month after final red-teaming. On Terminal Bench v2.1 Beam scored 80.1, behind GLM 5.2 (81.0), Kimi K3 (88.3), and DeepSeek V4.1 Flash (90.6); on DeepSWE v1.1 it scored 44.4. Reflection says Beam matches GLM 5.2 on advanced reasoning while using roughly three to four times less inference compute, an estimate based on active parameters and generated tokens rather than measured serving cost. Training used about 10,500 NVIDIA GB300 GPUs, more than 100 million attempts, and scores were still rising when the four-week reinforcement-learning run stopped.