Back to Home

Beam: Reflection's 501B open-weight model

Article URL: https://reflection.ai/blog/introducing-beam Comments URL: https://news.ycombinator.com/item?id=49969183 Points: 171 # Comments: 50

t
tech4you AI
October 6, 202610 min read
Share

We are introducing Beam, Reflection’s first open-weight model. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.

Figure 2: Beam demonstrates frontier-level inference efficiency, both when measured in terms of FLOPS and token count across DeepSWE, Humanity’s Last Exam (HLE), and Terminal Bench 2.1. We used data from Artificial Analysis and DataCurve, estimating generation forward-pass compute as FLOPs ≈ 2 × active parameter count × mean generated tokens per attempt, counting each multiply-add as two operations. Generated tokens include both reasoning and the final answer. For mixture-of-experts models, we used the parameters activated per token rather than the total model size. These estimates exclude prompt prefill, context-dependent attention operations, and serving overhead, so they represent an approximate compute comparison rather than measured inference cost. We use Artificial Analysis and DataCurve as sources for other model’s evals.
Figure 3: Terminal-Bench 2.1, HLE, and DeepSWE scores as a function of cumulative RL rollouts during training of Beam’s reasoning expert, accounting for 80M of the over 100M rollouts generated across the full RL campaign. For comparison, Inkling was trained on 30M rollouts and MiMo on 753K.
Figure 4: Stable learning continues as staleness builds up over time. The top plot shows the oldest sample in the batch, with the bottom plot demonstrating stable numerics. Even when training Beam with one-day staleness—107 weight versions behind the current policy—the numerics remain stable.
Figure 5: DeepSWE scores during part of the RL run. Each point corresponds to a different reasoning effort. The Pareto frontier moves in two phases. First, it contracts as the policy learns to be more token efficient. Then, the higher reasoning efforts expand outwards to achieve high performance.

Scaling reinforcement learning environments

Frontier-scale reinforcement learning requires a large volume of difficult, high-quality tasks. We built a pool of nearly one million environments, primarily through synthetic data pipelines, supplemented by proprietary vendor data and open-source sources.

We relied heavily on an iterative curation process. First, we synthesized or sourced environments across a broad set of domains including software engineering, terminal use, competitive coding, STEM, web search, tool use, and general knowledge work. Second, we heavily filtered tasks for difficulty (ensuring they were neither consistently solvable, nor impossible for the model) and quality (e.g., not underspecified, misleading, guessable, hackable, or otherwise broken or noisy). Third, we tested the tasks through RL, which allowed us to identify further quality or difficulty issues and inform the next iteration of sourcing and filtering.

Throughout Beam’s development, we found that compromises in data quality led to capability plateaus and other training issues. Systematic improvements to task quality were essential to sustaining capability gains throughout the run, which ended with no sign of saturation.

Frontier RL infrastructure

High-compute agentic RL requires generating rollouts, executing tools, evaluating outcomes, and updating the model at scale. We built an asynchronous platform that lets these processes run independently while coordinating the flow of experience and model updates.

During Beam’s training, we sustained an average of 110K concurrent rollouts. Seven capabilities made this practical:

Fully asynchronous execution: Agents generate rollouts while the trainer learns and publishes new model versions. Each token is tagged with the version that produced it, allowing the training algorithm to account for policy staleness as completed rollouts flow into training.

Flexible compute allocation: We adjusted the balance between inference and training as the workload evolved, operating at inference-to-training GPU ratios from 3.9:1 to 5.4:1. We also resized the trainer across five GPU mesh configurations within the same training lineage without losing training state.

Fast model updates: New weights reached the inference fleet in a median of approximately 12 seconds. Hierarchical distribution transfers weights across racks over RoCE, then shares them locally over NVLink. Compared with every replica pulling weights directly, this reduced cross-rack traffic by 75% and made fleet-wide adoption of new weights 2.2× faster.

Resilience to inference failures: During the run, 71 inference incidents were handled without terminating the training job. Inference capacity recovered in a median of eight minutes, with lost capacity accounting for just 0.02% of elapsed serving GPU-minutes.

Environments at scale: We supported up to 170K concurrent sandboxes during the run. Across the platform, we processed more than one billion sandbox creation requests, spanning over 20 clusters, two clouds, and four regions. 90% of new sandboxes were ready in under 10 seconds.

Efficient Trainer Packing: Dynamic packing kept training batches 99.99% full on average, holding per-GPU trainer throughput within 1.5% as mean rollout length grew almost 70%.

Observability and reward integrity: Per-token records enabled numerical consistency checks between training and inference at every step. Independent judges re-screened passing solutions for verifier exploits, while replayable records made rewards and their use in training inspectable.

Together, these capabilities enabled us to train on longer interactions and more demanding environments while maintaining throughput, recovering from failures, and checking the integrity of the learning process.

Figure 6: Beam’s pretraining recipe scales predictably across four orders of magnitude in compute. It achieves Pareto-optimal loss – compute performance on decontaminated code and web validation data among all comparable open base models.
Figure 7: Beam’s expert utilization is near-uniform. The busiest expert’s load, averaged across MoE layers, reaches just 1.04× at pretraining completion.
Figure 8: Residual stream RMS remains bounded throughout pretraining across all 52 layers of Beam, with smooth, depth-dependent trajectories and no sustained activation growth.

Quality-centric data curation

Beam was pretrained on 23.8 trillion diverse high-quality tokens from the web, public sources, and proprietary licensed datasets. Our data pipeline was designed to give Beam a foundation for downstream agentic coding: source code, technical explanations, and mathematical and scientific knowledge, preserved through every stage of curation. We train on almost all publicly accessible and unrestrictively-licensed code and code documentation on the web.

We trained our own quality classifiers for web, code, and STEM content, divided data into fine-grained quality tiers, and weighed training toward stronger material. After extensive scientific iteration, we optimized both precision and recall of data curation significantly beyond conventional web filters used in state-of-the-art OSS data frameworks. On one hand, about 95% of raw Internet tokens are eliminated through parsing, deduplication, and curation. On the other hand, we found that conventional techniques would have missed roughly 1.8 trillion high-quality tokens we retain, including 87% of our curated web-code tokens.

Code modeling requires its own curation for the highest performance. For each language, we applied individually tuned filters, removed low-quality autogenerated and unlearnable content, and trained classifiers to identify corrupted content or code that may hurt training stability. Overrepresented languages and file types are rebalanced to further broaden exposure.

We also developed a high-throughput pipeline for processing PDF artifacts, to ensure the base Beam has knowledge across a wide range of STEM topics. It integrates a vision-language OCR model with in-house quality classifiers and artifact detectors that catch malformed reconstructions, distributed across thousands of GPUs to process petabytes of technical data.

We repeated code and technical content multiple times to increase the model’s exposure over the course of its training horizon. This required careful attention to fuzzy deduplication, packing algorithms, and the science of overtraining, to ensure every repeated data source helps rather than hurts generalization.

Frontier-grade pretraining infrastructure

Beam was pretrained end-to-end in under four weeks on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs. To achieve the reliability, performance, and development velocity required to train Beam at scale, we built nearly the entire infrastructure stack in-house. This included a novel topology-aware, Kubernetes-based scheduler across our clusters; an internal node lifecycle system with continuous health monitoring and alerts; and a silent data corruption (SDC) detection system capable of semi-autonomous rewinds and restarts. Together, these systems gave us the performance and operational control required to train our own frontier model efficiently and with greater stability.

As a result of extensive investments in the training recipe stability, infrastructure, and data quality, the overall pretraining run finished with an extremely smooth trajectory. We executed nine semi-automatic rewinds throughout, attributed either to non-deterministic gradient norm spikes or to suspected SDCs. In addition, the run’s goodput (the share of wall-clock time spent on training steps retained in the final model) reached 92.3% towards the end thanks to improvements in checkpointing, fault detection, and node health management.

Figure 9: Beam training loss over the course of its pretraining run. We observed no instabilities or large irrecoverable spikes.
Figure 10: Alignment & safety RL predictably shapes Beam’s behavior across non-verifiable domains. From the pre-RL checkpoint, we are able to accurately forecast (dashed line) improvement in training reward (dots).

Originally published on Hacker News (Best)

Related Articles