Table of Contents
Reflection has introduced Beam, a 501-billion-parameter open-weight Mixture-of-Experts model built for coding, reasoning, and agentic workloads.
The U.S. AI lab said Beam activates 23 billion parameters per token, giving the model a much smaller inference footprint than its total parameter count suggests. The release is a Western open-weight alternative in a market where Chinese labs such as Qwen, GLM, Kimi, and DeepSeek have set much of the recent pace in coding and agentic benchmarks.
Beam is not fully public yet. Reflection said the model is still undergoing final red-teaming and evaluations, with weights, a technical report, a model card, documentation, and developer artifacts planned for release later this month under an Apache 2.0 license. Early access is available through the company's waitlist.

What Beam Is
Beam is a sparse MoE model, meaning only part of the full model is active for a given token. That design lets Reflection build a system with 501 billion total parameters while using 23 billion active parameters during inference.
The active-parameter count shapes the compute required to run the model. A model can carry a large pool of expert weights while routing each token through a smaller subset of them. Reflection's pitch is that Beam can deliver strong coding and agentic performance without the per-token inference cost of much larger dense or high-activation open models.
Beam was trained end-to-end from scratch. Reflection pretrained the model on 23.8 trillion curated tokens from the web, public sources, and proprietary licensed datasets, with an emphasis on code, technical documentation, STEM material, and agentic coding data. The pretraining run finished in under four weeks on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs, according to the company.
The model also went through a midtraining stage designed to prepare it for reinforcement learning and long-horizon tool use. Reflection said that stage extends Beam's effective context length to 1 million tokens, while the reinforcement learning run used a maximum context length of 256,000 tokens.
Reflection Claims Coding and Agentic Strength
Reflection is aiming Beam at coding, terminal work, reasoning, and agentic workflows. In the company's benchmark table, Beam posts company-reported scores of 44.4 on DeepSWE v1.1, 80.1 on Terminal Bench v2.1, 80.9 on SWEBench Verified, 77.2 on SWE Bench Pro v2-Hard, and 65.5 on SWE Bench Pro v1.
Beam is competitive with larger open models such as GLM 5.2 and is approaching Qwen 3.8-Max on coding and agentic tasks. Kimi K3 remains ahead on raw capability, while Beam's main advantage is inference efficiency rather than simply topping every benchmark.
Beam is being presented as a reliable workhorse for coding and agentic workloads rather than as the top scorer on every benchmark. Its comparison centers on capability per unit of inference compute.

Reflection says Beam reaches scores comparable to GLM-5.2 on advanced reasoning benchmarks while using 3–4× less inference compute.
The company estimated generation compute using active parameter count and mean generated tokens, while noting that the estimate excludes prompt prefill, context-dependent attention operations, and serving overhead.
A Large Reinforcement Learning Run
Reflection put much of its announcement on the scale of Beam's reinforcement learning. The company said it used 10.5K NVIDIA GB300 GPUs over four weeks, generating more than 100 million rollouts across coding, agentic, terminal, STEM, tool-use, and general-knowledge environments.
The run was one of the largest-scale reinforcement learning efforts conducted by an open lab to date. Training and grading used roughly 1.3 billion sandboxes, with an average of 110,000 concurrent rollouts and up to 170,000 concurrent sandboxes during the run.

The RL system was designed around long interactions, tool execution, environment feedback, and asynchronous training. Reflection said it trained Beam with asynchronous policy gradients, then developed methods to reduce instability from policy staleness when rollouts were generated by older model checkpoints.
That setup suits agentic systems, where the model does more than produce short answers. It may need to use tools, inspect files, write code, debug failures, search for information, and continue across multiple steps.
Beam's RL training produced gains that generalized beyond the specific tasks in the training mixture, including browsing behavior even when browsing tasks were not included in part of the RL mix.
Why Beam Matters for Open-Weight AI
Open-weight AI competition is increasingly defined by both capability and deployability. Developers and enterprises may want frontier-style coding and agentic performance, but running the largest models can be expensive, operationally difficult, or dependent on closed APIs.
Reflection is trying to occupy the middle ground with a large open-weight model that pairs a high total parameter count with a lower active inference footprint, long-context support, and planned Apache 2.0 licensing. The company also said it will release quantized FP8 and NVFP4 builds, which could make Beam easier to run across more infrastructure once the weights are public.

The release also gives Reflection a clearer position in the open-model race.
Beyond the announcement itself, the company is making a case that Western open-weight labs can build frontier-scale systems with their own data pipelines, training infrastructure, reinforcement learning environments, and safety processes.
→ Read Reflection's Beam announcement and join the early-access waitlist