Reflection’s Beam trails top open models on coding tests but claims lower inference compute
Reflection AI’s Beam, a 501B open-weight coding model, trails top open models but claims lower inference compute.
Reflection AI announced Beam, a 501-billion-parameter open-weight mixture-of-experts model for coding and agents, with 23 billion parameters active per token. It plans to release the weights under Apache 2.0 later this month after final red-teaming. On Terminal Bench v2.1 Beam scored 80.1, behind GLM 5.2 (81.0), Kimi K3 (88.3), and DeepSeek V4.1 Flash (90.6); on DeepSWE v1.1 it scored 44.4. Reflection says Beam matches GLM 5.2 on advanced reasoning while using roughly three to four times less inference compute, an estimate based on active parameters and generated tokens rather than measured serving cost. Training used about 10,500 NVIDIA GB300 GPUs, more than 100 million attempts, and scores were still rising when the four-week reinforcement-learning run stopped.
- Beam is a 501-billion-parameter mixture-of-experts model activating 23 billion parameters per token.
- Apache 2.0 weights are planned later this month after final red-teaming.
- Terminal Bench v2.1 score was 80.1, trailing DeepSeek V4.1 Flash at 90.6.
- Reflection claims reasoning comparable to GLM 5.2 at three to four times less inference compute.
- Training used about 10,500 NVIDIA GB300 GPUs; scores were still climbing when the run ended.
Full article444 words · extracted from helpnetsecurity.com · click to collapse
Reflection AI has built Beam, a 501-billion-parameter open-weight model for coding and agent tasks, and plans to publish the weights under an Apache 2.0 license later this month. Open-weight means developers can download the trained model and run it on their own hardware. Beam uses a mixture-of-experts design, so 23 billion of its parameters fire for any given token, which keeps the cost of each answer down.

Beam does not lead the open field on raw scores. On Terminal Bench v2.1, a test of command-line work, Beam scored 80.1, against 81.0 for GLM 5.2, 88.3 for Kimi K3 and 90.6 for DeepSeek V4.1 Flash. On DeepSWE v1.1 it scored 44.4, just above GLM 5.2’s 44.0 and 29.8 points behind DeepSeek V4.1 Flash. Reflection says Beam scores comparably to GLM 5.2 on advanced reasoning benchmarks while using three to four times less inference compute.
Treat that compute figure as a rough guide. Reflection estimated it with a formula, twice the active parameter count times the average number of tokens generated, and calls the result an approximate comparison of compute that does not measure inference cost. The formula leaves out prompt processing, attention costs and serving overhead. Scores for competing models came from Artificial Analysis and DataCurve. If you run agents around the clock, your real bill depends on your prompts and hardware, and that test is yours to run.
The training run had not leveled off
Reflection spent four weeks on reinforcement learning, the stage where a model improves by attempting tasks and getting graded. The run used about 10,500 NVIDIA GB300 GPUs, generated more than 100 million attempts, drew on nearly one million training environments and used roughly 1.3 billion sandboxes for training and grading. The company says scores were still climbing when the run ended.
Some of what Beam learned spread beyond its training tasks. Its browsing scores rose during a phase whose task mix included no browsing, and when given web access it began searching for and querying other large language models and calling OCR services to read documents. That is useful in an agent. It is also behavior worth logging before you give Beam network access.
Safety results come later
Beam is still in final red-teaming and evaluation. Reflection trained a separate safety and alignment model from the same starting checkpoint, then combined the two through distillation, a process that trains one model to imitate others. The company plans to publish its safety results in a technical report released alongside the weights, and to open-source the safety tests it built internally. Early access runs through a waitlist for a select group of users.

Webinar: Closing the accountability gap in AI-assisted delivery