Beam collected 302 points and 78 comments on Hacker News before developers could download a single weight. Reflection calls the 501-billion-parameter model open-weight, but its launch post sends readers to an early-access waitlist and says the weights, technical report, model card, and developer materials will arrive later in October. That timing is the part developers should hold onto: the company has disclosed a remarkable amount about the training run, while the evidence needed to reproduce its results is still pending.
The model itself is substantial. Beam is a sparse mixture-of-experts system with 501 billion total parameters and 23 billion active for each token. Reflection says it was pretrained on 23.8 trillion tokens, then put through more than 100 million reinforcement-learning rollouts. The company's announcement targets coding, reasoning, tool use, and longer agent tasks. Beam is text-only, with an effective context length of one million tokens after midtraining.
Those specifications explain the immediate interest. They do not yet tell a team whether Beam belongs in a production stack. Reflection's benchmark tables are company results, and TechCrunch reports that they have not been independently verified. Until the files and evaluation details ship, Beam is a detailed preview of an open-weight release rather than a release that outside developers can inspect.
A smaller active model inside a very large one
The 501 billion figure describes all of Beam's parameters, not the amount used for every token. Its mixture-of-experts architecture routes each token through a subset totaling 23 billion active parameters. That distinction affects memory, throughput, and serving cost, which is why Reflection emphasizes active parameters when comparing Beam with other sparse models in its technical overview. The full checkpoint will still be large to store and distribute even when inference activates a much smaller slice at once.
Reflection's own tables place Beam among current open-weight systems without claiming a clean sweep. Beam scores 80.1 on Terminal-Bench 2.1, compared with 88.3 for Kimi K3 and 90.6 for DeepSeek V4.1 Flash. On SWE-Bench Pro v2-Hard, Beam records 77.2, behind GLM 5.3 at 84.3 and Kimi K3 at 88.2. It reaches 80.9 on SWE-Bench Verified, though the table does not report scores for several of the larger comparison models on that test. These company-reported results support a case for competitiveness, not across-the-board leadership.
The pitch rests instead on the work done per unit of inference compute. Reflection says Beam can match GLM 5.2 on advanced reasoning tests while using three to four times less inference compute. Its estimate multiplies active parameters by the mean number of generated tokens. The methodology note excludes prompt prefill, context-dependent attention, and serving overhead, so the chart is an approximation rather than a measured bill from a deployed service. Hardware utilization, batching, quantization, memory traffic, and the chosen reasoning setting can all change the price a user ultimately pays.
That caveat does not make the comparison useless. It defines what the number can answer. Reflection says it trained Beam to spend fewer generated tokens on some successful solutions, then allow longer traces when harder tasks benefit from them. Beam has a reasoning-effort control for that tradeoff, according to the launch post. An independent evaluation should therefore compare task success, tokens, latency, and actual serving cost at matched settings, rather than treating one benchmark score as a complete efficiency result.
Ten thousand GPUs and 1.3 billion sandboxes
Beam's most unusual numbers concern the machinery behind the model. Reflection says the reinforcement-learning phase used 10,500 Nvidia GB300 GPUs for four weeks, produced more than 100 million rollouts, and used approximately 1.3 billion sandboxes for training and grading. The system averaged 110,000 concurrent rollouts and supported as many as 170,000 concurrent sandboxes. According to the company's account, new weights reached the inference fleet in a median of about 12 seconds while agents continued generating experience.
That asynchronous design creates an awkward problem: a long rollout may contain tokens produced by several model checkpoints, and completed work can reach the trainer after the policy has changed. Reflection says its training remained stable even when examples were more than a day old and 107 weight versions behind. It tagged tokens with the version that generated them and developed methods to reduce mismatch between training and inference. The reported solution will interest teams working on agent training because keeping thousands of tool-using jobs productive is a systems problem as much as a model problem.
The infrastructure also had to tolerate ordinary failure at an uncommon scale. Reflection reports 71 inference incidents during the run, with capacity recovering in a median of eight minutes. It says dynamic packing kept training batches 99.99% full on average, even as mean rollout length grew by almost 70 percent. These figures come from Reflection's own operational summary. The promised technical report will need to explain enough of the setup for readers to judge how the measurements were defined.
Pretraining was a separate four-week run across 6,144 GB300 GPUs. Reflection says it completed nine semi-automatic rewinds after gradient spikes or suspected silent data corruption and reached 92.3 percent goodput near the end. The company also says its filters discarded about 95 percent of raw internet tokens while retaining 1.8 trillion tokens that conventional filtering would have missed. The post describes public web material and proprietary licensed data, but it does not publish a full training-data inventory.
What has to ship in October
Reflection says Beam's weights will ship under Apache 2.0, accompanied by documentation and a stack for running, evaluating, and fine-tuning the model. If that package arrives as described, developers will be able to check memory requirements, quantization behavior, tokenizer details, long-context performance, and the benchmark harnesses themselves. The announced license and October schedule are more useful than a vague commitment, though they remain future deliverables on launch day.
The same gap applies to safety. Beam is undergoing final red-team work and evaluation. Reflection says it trained a separate safety and alignment model, merged that work with the capability model through multi-teacher on-policy distillation, and plans to publish safety results in the technical report. It also plans to release internally developed safety evaluations. Until those materials appear, outside reviewers cannot examine the policy, test refusal behavior, or compare agent actions under adversarial prompts using the company's stated evaluation setup.
The Hacker News discussion reflects that unfinished state. Commenters focused on the absent weights, comparisons with models that are already downloadable, and questions about the evaluation claims. That discussion is a measure of developer attention, not proof that Beam succeeds or fails. It does show what the launch numbers could not settle: readers want an artifact they can run.
Reflection has the funding and compute to make the release consequential if it follows through. TechCrunch says the two-year-old company has raised roughly $4.7 billion and signed compute agreements worth more than $7 billion with SpaceX and Nebius for access to Nvidia GB300 hardware through 2029. The company is pitching Beam to developers, enterprises, and governments that want to customize and operate models on their own infrastructure. For those buyers, control over weights matters only when deployment requirements and license terms can be checked in practice.
October's promised release now has a clear test. Watch for the checkpoint under the announced Apache 2.0 license, a model card that states hardware and known limits, safety results with reproducible methods, and evaluation code that resolves the missing details behind the efficiency chart. If those pieces land together, Beam can be judged as an open-weight model. Until then, its strongest public artifact is a precise account of an enormous training run, attached to a waitlist.