Beam: Reflection Launches an Open 501B MoE Betting on Massive RL
With 501 billion parameters, 23 billion active and 100 million reinforcement learning rollouts, Reflection wants to prove the West can still matter in open-weight.

In brief
Reflection unveils Beam, its first open-weight model: a Mixture-of-Experts with 501B parameters (23B active) built for code and agents, with weights shipping this month under Apache 2.0 license. It doesn't beat the top Chinese open models on raw capability, but claims 3 to 4 times better inference efficiency at comparable capability. The post mostly details an extraordinary RL campaign and the infrastructure that made it possible.
🍺 Bar-stool version
Reflection is releasing an open model that isn't the top of the class, and it owns that: its pitch is that it gets nearly as far while burning three to four times less compute. It's a bit like the car that doesn't win the Grand Prix but drives Paris to Marseille on half a tank. To pull that off, they ran 10,500 GPUs for a month just to teach it to think better, which is a pretty peculiar definition of frugality. If the weights really do ship under Apache 2.0, companies wanting a powerful, self-hostable, non-Chinese code model will finally have a serious option.
Key takeaways
- 1
Beam is a sparse MoE with 501 billion parameters, 23 billion active per token, pre-trained on 23.8 trillion tokens, with effective context extended to 1M tokens.
- 2
The RL phase used 10,500 NVIDIA GB300 GPUs for four weeks, over 100 million rollouts, roughly 1.3 billion sandboxes and nearly a million environments.
- 3
On agentic coding benchmarks, Beam scores 80.9 on SWE-bench Verified and 80.1 on Terminal Bench 2.1, behind GLM 5.3, Kimi K3, Qwen 3.8 Max and DeepSeek V4.1 Flash on the latter.
- 4
Reflection claims reasoning scores comparable to GLM 5.2 with 3 to 4 times less inference compute, based on a FLOPs estimate it itself calls approximate.
- 5
Capabilities showed no plateau as RL compute increased, and web navigation skills emerged without any browsing tasks in training.
- 6
Weights, the technical report and the model card will be published this month under Apache 2.0 license; the model is currently in early access via waitlist.
A workhorse model rather than a champion
Reflection presents Beam as its first open-weight model, built for code, reasoning and agentic workloads. The architecture is a Mixture-of-Experts: 501 billion parameters total, but only 23 billion engaged per token.
The positioning is deliberate. Beam claims to be competitive with GLM 5.2 and close to Qwen 3.8-Max on code and agentic tasks, while acknowledging Kimi K3 remains ahead in raw capability. The pitch is inference efficiency.
The published numbers confirm this nuanced picture: 80.9 on SWE-bench Verified, 77.2 on SWE Bench Pro v2-Hard, but 44.4 on DeepSWE versus 68.0 for Kimi K3 and 74.2 for DeepSeek V4.1 Flash. The efficiency comparison relies on an estimate (2 × active parameters × generated tokens) that excludes prefill and serving costs.
RL as the main scaling axis
The core of the post focuses on reinforcement learning. Reflection mobilized 10,500 GB300 GPUs for four weeks to generate over 100 million rollouts, with contexts up to 256K tokens. For comparison, the lab notes that Inkling was trained on 30 million rollouts and MiMo on 753,000.
The lab claims scores on Terminal-Bench, HLE and DeepSWE kept improving with RL compute, with no visible plateau. It used fully asynchronous policy gradients, with new algorithms to stay stable even when learning from interactions over a day old, i.e. 107 weight versions behind.
A controllable length penalty first taught the model to solve tasks with fewer tokens, before agentic capabilities justified longer responses. The user retains control via a reasoning effort parameter.
Generalization and environments
Trained on reasoning, software engineering and terminal tasks, Beam improved at web navigation despite never having browsing tasks. Given web access, it learned on its own to query other LLMs and use OCR APIs to read documents.
This RL relies on nearly a million environments, mostly synthetic, supplemented by vendor and open source data. They were filtered to be neither trivial nor impossible, nor hackable. According to Reflection, every compromise on task quality led to capability plateaus.
Infrastructure built in-house
The RL platform supported an average of 110,000 concurrent rollouts and up to 170,000 concurrent sandboxes, across more than 20 clusters, two clouds and four regions. New weights reached the inference fleet in about 12 seconds median, and 71 inference incidents were absorbed without halting training.
Pre-training ran in under four weeks on 6,144 GB300 NVL72 GPUs, with an in-house Kubernetes scheduler and a system for detecting silent data corruption. Result: only nine semi-automatic rewinds and 92.3% goodput by the end of the run.
Data, architecture and alignment
The 23.8 trillion tokens come from the web, public sources and proprietary licensed datasets. About 95% of raw tokens are discarded, but Reflection claims to recover 1.8 trillion quality tokens that standard filters would have dropped, 87% of which from its curated web code.
The architecture combines interleaved local and global attention, finely routed experts and several load balancing mechanisms inspired by DeepSeek. The most loaded expert never exceeds 1.04 times the average by the end of pre-training, across 52 layers.
On safety, a second model was trained specifically on alignment, then merged with the RL model via on-policy multi-teacher distillation. Reflection promises to publish its safety evaluations and open-source its tools.
“Beam advances the Western open-weight frontier and is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks.”
“We believe this is one of the largest scale RL runs conducted by any open lab to date.”
“Throughout Beam's development, we found that compromises in data quality led to capability plateaus and other training issues.”
Why it matters
Beam arrives in an open-weight landscape largely dominated by Chinese labs (Qwen, Kimi, GLM, DeepSeek), and Reflection explicitly frames it as advancing the 'Western' open frontier. That's the real stake: giving companies a self-hostable code model under Apache 2.0 without depending on a Chinese provider. The post is also unusually transparent about the mechanics of large-scale RL, infrastructure and data curation, making it valuable reading for anyone training models. Still, caution is warranted: the weights aren't published yet, benchmarks show Beam behind several open competitors on Terminal Bench and DeepSWE, and the central efficiency claim rests on a theoretical FLOPs estimate, not a measured inference cost. The claim of 'no plateau' progress is the most scientifically interesting, and the one most worth verifying once the technical report is available.
For you
Put it to work on your sources.
Free: a month of articles and three sources of your own. Pro: the whole archive and your sources, from €8/month.
For your team
The same machine, on your topics.
A space in your colours, your watch angles, your curators. Pilot open to three companies.
Read next
#agentsTodayCloudflare launches Clef-omni: a decision model that sees, hears, and decides
A week after its first open-weight models, Cloudflare adds audio and video, slashes Clef-flash's price by more than half, and speeds up Clef.
Source · Cloudflare Blog · Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash
#geminiTodayGemini agent: Google Cloud wants one agent for all your work
At Gemini at Work 2026, Thomas Kurian unveils a single, persistent, multi-model agent that even gets its own corporate email address.
Source · Google Cloud Blog · Gemini at Work 2026: Introducing Gemini agent
#privacyTodayAdam Schiff: Congress vs. AI, Caught Between Section 230 and Murphy's Law
The Democratic senator from California tears apart the voluntary pact signed at the White House and pushes for an "FDA for AI"—while admitting Congress has never known how to regulate tech.
Source · The Verge – Decoder · Sen. Adam Schiff on making Congress relevant in the AI age