Reflection has introduced Beam, its first planned open-weight model, with a focus on coding, reasoning and agentic workloads. The company describes a sparse mixture-of-experts architecture with 501 billion total parameters and 23 billion active parameters. It is presenting early performance results while final red-teaming and evaluations continue.
The weights are not part of the announcement's immediate release. Reflection says it will publish the weights, technical report, model card and developer artifacts later this month, and offers an early-access signup. For developers deciding whether to build with Beam, that makes this a preview to investigate rather than a model package already available for independent deployment.
A large model with a smaller active computation path
In a mixture-of-experts model, the total parameter count and the number used for a token are different measures. Reflection uses that distinction to frame Beam's appeal: a large overall model with a smaller active parameter count, trained to avoid spending unnecessary reasoning tokens.
The company reports pretraining on 23.8 trillion tokens drawn from web material and licensed datasets. Its announcement then describes a substantial reinforcement-learning campaign: more than 100 million rollouts, 10,500 NVIDIA GB300 GPUs and four weeks of training. These figures are Reflection's account of its training process, not a measurement conducted by Franklin.
Beam is text-only. Reflection's demonstrations nevertheless involve tasks with visual or external information, handled through text representations and tools. The distinction matters when judging the examples: using an OCR service or describing a game's appearance in text does not make the underlying model natively multimodal.
Benchmark claims need their measurement context
Reflection publishes results across coding, terminal work, reasoning, tool calling and search. It positions Beam as competitive with some larger open models while acknowledging that other frontier open models remain ahead on raw capability. That is a more limited claim than being the best model across all tasks.
Its efficiency comparison estimates generation compute from active parameters and mean generated tokens. Reflection explicitly excludes prompt prefill, context-dependent attention operations and serving overhead. Those exclusions prevent the chart from serving as a complete hosting-cost or latency comparison. A model that uses fewer estimated generation operations can still have different memory, infrastructure and request-handling requirements.
The announcement also gives users a reasoning-effort control. Lower settings favor shorter responses; higher settings allow longer reasoning on demanding tasks. Developers will need to test that tradeoff on their own workloads rather than assuming the highest setting is always the most economical choice.
Training agents means checking the environment too
Reflection describes a pool of nearly one million reinforcement-learning environments, spanning software engineering, terminal use, STEM, web search and other tasks. It says the curation process filtered environments for difficulty and quality, including misleading instructions, broken tasks and opportunities to exploit a verifier.
That concern connects to the problem examined by A2Z GameSpec-Bench: working software can still violate its specification. Reflection discusses the quality of the training tasks and rewards; the game benchmark checks whether generated output follows a designer's rules. Both focus on whether apparent success matches the intended task, rather than treating a runnable result as sufficient.
Reflection says its asynchronous infrastructure tags tokens with the model version that generated them, so learning can account for older rollouts. It also describes independent judges rechecking passing solutions for verifier exploits. These are relevant implementation details because a successful reward signal is only useful when it represents the behavior the developer wanted to teach.
What remains to evaluate after release
The forthcoming model card, technical report and weights should make it easier to inspect the conditions behind the preview. Developers still need the actual release terms, deployment requirements and task-level behavior before treating Beam as a production option.
Reflection's stated combination of coding capability and shorter reasoning is worth testing against a fixed task set. Keep the task, tools and success criteria consistent, then compare quality alongside actual latency and serving cost. The preview provides a reason to run that evaluation; it does not replace it.
