Beam: Reflection's 501B open-weight model
Article URL: https://reflection.ai/blog/introducing-beam Comments URL: https://news.ycombinator.com/item?id=49969183 Points: 171 # Comments: 50
We are introducing Beam, Reflection’s first open-weight model. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.





Scaling reinforcement learning environments
Frontier-scale reinforcement learning requires a large volume of difficult, high-quality tasks. We built a pool of nearly one million environments, primarily through synthetic data pipelines, supplemented by proprietary vendor data and open-source sources.
We relied heavily on an iterative curation process. First, we synthesized or sourced environments across a broad set of domains including software engineering, terminal use, competitive coding, STEM, web search, tool use, and general knowledge work. Second, we heavily filtered tasks for difficulty (ensuring they were neither consistently solvable, nor impossible for the model) and quality (e.g., not underspecified, misleading, guessable, hackable, or otherwise broken or noisy). Third, we tested the tasks through RL, which allowed us to identify further quality or difficulty issues and inform the next iteration of sourcing and filtering.
Throughout Beam’s development, we found that compromises in data quality led to capability plateaus and other training issues. Systematic improvements to task quality were essential to sustaining capability gains throughout the run, which ended with no sign of saturation.
Frontier RL infrastructure
High-compute agentic RL requires generating rollouts, executing tools, evaluating outcomes, and updating the model at scale. We built an asynchronous platform that lets these processes run independently while coordinating the flow of experience and model updates.
During Beam’s training, we sustained an average of 110K concurrent rollouts. Seven capabilities made this practical:
Fully asynchronous execution: Agents generate rollouts while the trainer learns and publishes new model versions. Each token is tagged with the version that produced it, allowing the training algorithm to account for policy staleness as completed rollouts flow into training.
Flexible compute allocation: We adjusted the balance between inference and training as the workload evolved, operating at inference-to-training GPU ratios from 3.9:1 to 5.4:1. We also resized the trainer across five GPU mesh configurations within the same training lineage without losing training state.
Fast model updates: New weights reached the inference fleet in a median of approximately 12 seconds. Hierarchical distribution transfers weights across racks over RoCE, then shares them locally over NVLink. Compared with every replica pulling weights directly, this reduced cross-rack traffic by 75% and made fleet-wide adoption of new weights 2.2× faster.
Resilience to inference failures: During the run, 71 inference incidents were handled without terminating the training job. Inference capacity recovered in a median of eight minutes, with lost capacity accounting for just 0.02% of elapsed serving GPU-minutes.
Environments at scale: We supported up to 170K concurrent sandboxes during the run. Across the platform, we processed more than one billion sandbox creation requests, spanning over 20 clusters, two clouds, and four regions. 90% of new sandboxes were ready in under 10 seconds.
Efficient Trainer Packing: Dynamic packing kept training batches 99.99% full on average, holding per-GPU trainer throughput within 1.5% as mean rollout length grew almost 70%.
Observability and reward integrity: Per-token records enabled numerical consistency checks between training and inference at every step. Independent judges re-screened passing solutions for verifier exploits, while replayable records made rewards and their use in training inspectable.
Together, these capabilities enabled us to train on longer interactions and more demanding environments while maintaining throughput, recovering from failures, and checking the integrity of the learning process.



Quality-centric data curation
Beam was pretrained on 23.8 trillion diverse high-quality tokens from the web, public sources, and proprietary licensed datasets. Our data pipeline was designed to give Beam a foundation for downstream agentic coding: source code, technical explanations, and mathematical and scientific knowledge, preserved through every stage of curation. We train on almost all publicly accessible and unrestrictively-licensed code and code documentation on the web.
We trained our own quality classifiers for web, code, and STEM content, divided data into fine-grained quality tiers, and weighed training toward stronger material. After extensive scientific iteration, we optimized both precision and recall of data curation significantly beyond conventional web filters used in state-of-the-art OSS data frameworks. On one hand, about 95% of raw Internet tokens are eliminated through parsing, deduplication, and curation. On the other hand, we found that conventional techniques would have missed roughly 1.8 trillion high-quality tokens we retain, including 87% of our curated web-code tokens.
Code modeling requires its own curation for the highest performance. For each language, we applied individually tuned filters, removed low-quality autogenerated and unlearnable content, and trained classifiers to identify corrupted content or code that may hurt training stability. Overrepresented languages and file types are rebalanced to further broaden exposure.
We also developed a high-throughput pipeline for processing PDF artifacts, to ensure the base Beam has knowledge across a wide range of STEM topics. It integrates a vision-language OCR model with in-house quality classifiers and artifact detectors that catch malformed reconstructions, distributed across thousands of GPUs to process petabytes of technical data.
We repeated code and technical content multiple times to increase the model’s exposure over the course of its training horizon. This required careful attention to fuzzy deduplication, packing algorithms, and the science of overtraining, to ensure every repeated data source helps rather than hurts generalization.
Frontier-grade pretraining infrastructure
Beam was pretrained end-to-end in under four weeks on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs. To achieve the reliability, performance, and development velocity required to train Beam at scale, we built nearly the entire infrastructure stack in-house. This included a novel topology-aware, Kubernetes-based scheduler across our clusters; an internal node lifecycle system with continuous health monitoring and alerts; and a silent data corruption (SDC) detection system capable of semi-autonomous rewinds and restarts. Together, these systems gave us the performance and operational control required to train our own frontier model efficiently and with greater stability.
As a result of extensive investments in the training recipe stability, infrastructure, and data quality, the overall pretraining run finished with an extremely smooth trajectory. We executed nine semi-automatic rewinds throughout, attributed either to non-deterministic gradient norm spikes or to suspected SDCs. In addition, the run’s goodput (the share of wall-clock time spent on training steps retained in the final model) reached 92.3% towards the end thanks to improvements in checkpointing, fault detection, and node health management.


Originally published on Hacker News (Best)
