Lilith.
⌕
Editorial illustration: Beam activates 23 billion parameters, but its open weights are still to come
Lilith illustration · editorial remix

Reflection has introduced Beam, its first open-weight model for coding, reasoning and agentic workloads. It is a sparse Mixture-of-Experts model with 501 billion parameters, activating 23 billion for each token. The weights, technical report, model card and developer artifacts are due later in October under an Apache 2.0 license.

A huge model engages only a fraction of its capacity at inference

Reflection says it pretrained Beam on 23.8 trillion tokens. The reinforcement learning phase then produced more than 100 million rollouts over 4 weeks on 10,500 NVIDIA GB300 GPUs. The company says it used about 1.3 billion sandboxes and one million coding, agentic and STEM environments.

The model is currently in early access while final red-teaming and evaluations continue. Its open weights were therefore not available to download at announcement time. Open-weight describes the promised distribution model, not current availability.

The smaller active core targets operating cost, not file size

For operators, the 23 billion active parameters matter because they shape the computation needed to generate each token. All 501 billion parameters still affect storage, loading and the hardware footprint of serving the model. Sparse architecture reduces part of the bill, but it does not turn a very large model into workstation software.

Reflection is targeting enterprises and sovereign institutions that want to run and adapt a model on their own infrastructure. An Apache 2.0 release could support that goal. Practical value will still depend on documentation, inference framework support and clear hardware profiles.

Reflection itself calculates the claimed 3 to 4 times saving

The company says Beam matches GLM-5.2 on reasoning benchmarks while using 3 to 4 times less inference compute. The comparison estimates FLOPs from active parameter count and generated tokens. It excludes prompt prefill, context-dependent attention operations and serving overhead.

That makes it a useful technical estimate rather than a measured production price or latency result. Independent teams have not yet been able to verify either the benchmarks or the operating advantage because the weights were not public at launch.

The October release will show whether benchmarks become a deployable model

The first decisive signal will be the actual release of the weights, model card and technical report under the promised license. That will allow outsiders to measure throughput, latency, memory requirements and quality beyond the company's evaluation suite.

The second signal will be support in common open-source libraries and fine-tuning without a proprietary Reflection layer. Those details will show whether Beam is an open alternative or, for now, an impressively calculated promise.

Lilith's verdict

Reflection has shown an engine with 501 billion parts and says only 23 billion engage while it runs. Beam starts racing only when the hood opens and an outsider measures the fuel bill.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗