Technical research with cited sources. Original measurements are identified in the article.

This article was researched and written with AI assistance. Editorial responsibility: Julian Dominic Altmann. How this site is made

Published: October 6, 2026 Updated: October 6, 2026

About the author

Research checked: October 6, 2026. Reflection AI announced Beam on October 5, 2026 as a 501-billion-parameter sparse Mixture-of-Experts model with 23 billion parameters active per token, aimed at coding, reasoning, and agentic workloads. The most important limitation is easy to miss in launch coverage: the weights are not public yet. Reflection says the weights, model card, technical report, and developer artifacts will arrive later in October under an Apache 2.0 license. [S01][S03]

Beam looks promising because it targets strong agent and coding performance with only 23B parameters active per token, but the current performance picture is still dominated by Reflection’s own launch evaluations. Local deployment, independent reproduction, and meaningful Apple-silicon testing all have to wait for the weights. [S01][S12]

Beam specifications at a glance

ItemStatus on October 6, 2026
DeveloperReflection AI
AnnouncedOctober 5, 2026
ArchitectureSparse Mixture of Experts
Total parameters501B
Active parameters23B per token
Active shareabout 4.59%
Layers52
Pretraining data23.8T tokens
Effective trained context1M tokens
RL scalemore than 100M rollouts
RL hardware10.5K NVIDIA GB300 GPUs for four weeks
Pretraining hardware6,144 NVIDIA GB300 NVL72 GPUs; under four weeks
Modalitydescribed at launch as text-focused / text-only
Public weightsnot yet released
Planned licenseApache 2.0
Public API priceno authoritative public per-token rate found

Primary source: Reflection AI [S01]. Independent launch confirmation: Reuters [S03] and Axios [S05].

Why Beam is more interesting than the “501B” headline

A 501B model normally sounds prohibitively expensive. Beam’s design changes the inference equation: it is sparse, so only 23B parameters participate in the forward pass for each token. That is roughly 4.59% of the total parameter count. [S01]

That is a compute advantage, not a storage shortcut. The full expert set still has to be accessible to the runtime. A local machine therefore cannot treat Beam like an ordinary 23B model just because 23B parameters are active at once.

This distinction is central to understanding Reflection’s pitch. The company does not claim Beam is the absolute strongest open model on every benchmark. It explicitly says Kimi K3 remains ahead in raw capability and frames Beam around inference efficiency. [S01] Axios similarly describes Beam as competitive on selected tasks while emphasizing Reflection’s claim of GLM-5.2-like reasoning at lower inference compute. [S05]

Architecture: a large expert pool with a small active path

Reflection says Beam combines interleaved local and global attention, fine-grained routed experts, a controlled residual stream, and several balancing techniques. Its routing recipe builds on auxiliary-loss-free load balancing and adds cosine decay to expert-bias updates plus sequence-level balancing. The company reports that the busiest expert averaged just 1.04× the mean load by the end of pretraining. [S01]

The model has 52 layers. Reflection also describes depth-aware residual scaling, SandwichNorm, elementwise attention gating, and FP32 residual accumulation as part of its stability recipe. [S01]

These details matter because high-compute reinforcement learning is harder to sustain when expert routing or residual magnitudes become unstable. Beam’s architecture is clearly designed around post-training at scale, not just pretraining efficiency.

Beam as a 501B MoE with a 23B-active path.

Training scale: 23.8T tokens and more than 100M RL rollouts

Reflection says Beam was pretrained on 23.8 trillion curated tokens from the web, public sources, and proprietary licensed datasets. The company says roughly 95% of raw internet tokens were removed through parsing, deduplication, and quality filtering. [S01]

The pretraining run reportedly completed in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs. Reflection then ran reinforcement learning for four weeks on 10,500 GB300 GPUs, generating more than 100 million rollouts across roughly 1.3 billion sandboxes and about one million coding, agentic, and STEM environments. [S01]

Those numbers are unusual enough to be a major part of the story, but they are still vendor disclosures. The pending technical report will be important for judging the exact training mixture, eval contamination controls, reward design, safety methodology, and reproducibility.

Beam training scale: 23.8T tokens and over 100M RL rollouts.

The 1M context claim needs a serving caveat

Reflection states that midtraining extended Beam’s effective context length to 1 million tokens. It also says the large RL campaign used rollouts with a maximum context length of 256K. [S01]

Independent technical write-ups that reviewed Reflection’s beta developer documentation on October 5–6 report a current API budget of roughly 262,144 combined tokens, with up to about 131,072 output tokens. Those limits are described as beta behavior and may change. [S11][S12][S15]

The clean way to publish the spec is therefore:

  • 1M tokens: the official effective-context training claim.
  • about 256K/262K: the currently reported beta serving window.
  • about 128K/131K output: a reported beta cap, not necessarily the final product limit.

This avoids turning a training capability into an incorrect API guarantee.

Beam benchmarks: competitive, but not independently established yet

Reflection published a wide comparison against Inkling, Nemotron 3 Ultra, GLM-5.2, GLM-5.3, Kimi K3, Qwen3.8-Max, and DeepSeek V4.1 Flash. [S01]

Selected launch-table results:

BenchmarkBeamGLM-5.2GLM-5.3Kimi K3Qwen3.8-MaxDeepSeek V4.1 Flash
DeepSWE v1.144.444.061.068.051.074.2
SWE-bench Pro v165.562.1NRNR67.7NR
Terminal-Bench 2.180.181.088.288.386.690.6
HLE, no tools36.240.542.346.943.639.1
GPQA Diamond90.591.291.793.592.690.9
AutomationBench public37.026.248.246.739.854.8
AA-LCR79.378.379.788.780.384.0
IFBench79.773.3NRNR82.8NR

NR means Reflection did not report a comparison value in that row. [S01]

What the table actually says

Beam is competitive, but it is not the top model across the board. On Terminal-Bench 2.1, Beam’s 80.1 is narrowly below GLM-5.2 and materially below GLM-5.3, Kimi K3, Qwen3.8-Max, and DeepSeek V4.1 Flash. On DeepSWE v1.1, the gap to the strongest models is larger. [S01]

The more credible Beam thesis is therefore performance per unit of active inference compute, not blanket benchmark leadership.

Why benchmark methodology matters here

SWE-bench Verified is a 500-instance human-filtered software-engineering set, and its maintainers explicitly warn that results from different agent versions and harness revisions are not automatically comparable. [S22]

Terminal-Bench 2.1 is itself a revised benchmark that fixed 28 of the 89 tasks from version 2.0. [S25] AutomationBench distinguishes between its public tasks and a separate held-out official set, so public local scores need not match private leaderboard results one-for-one. [S23]

AA-LCR tests reasoning across long document sets averaging roughly 100K input tokens; it is not a general intelligence score. [S38] IFBench measures precise instruction following under verifiable constraints rather than software-engineering performance. [S39]

Most importantly, independent Beam evaluation is still thin. AIEvals reported no independent Beam result in the external leaderboards it checked as of the launch window. [S12] During this review, Beam also had not yet become a normal, reproducible open-weight entry in the same way as models with downloadable checkpoints.

Beam benchmark results with independent caveats.

Beam versus the open-weight field

GLM-5.2 and GLM-5.3

Z.ai’s GLM-5.2 already supports a 1M context and has downloadable weights. GLM-5.3 uses the same base-model foundation with stronger post-training and better complex coding performance. [S30][S31]

That makes the comparison asymmetric today: Beam may eventually offer better efficiency, but GLM can already be deployed and independently tested.

Qwen3.8-Max

Alibaba describes Qwen3.8-Max as a 2.4-trillion-parameter multimodal model with up to a 1M-token context. [S29] Reflection says Beam approaches it on coding and agentic tasks while operating at a much smaller active footprint. [S01]

That is a compelling efficiency story, but not yet a deployment advantage until Beam’s weights and runtime ecosystem exist.

Kimi K3

Moonshot AI positions Kimi K3 as an open-weight native multimodal agentic model. [S32] Reflection explicitly says Kimi K3 remains ahead on raw capability. [S01] Anyone optimizing purely for maximum quality should therefore avoid assuming Beam is a straight upgrade.

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is a 552B MoE with an asymmetric design that activates 8B parameters on input and 16B on output. It is already available via DeepSeek’s API. [S33][S40]

In Reflection’s own comparison table, V4.1 Flash leads Beam by a large margin on Terminal-Bench 2.1 and DeepSWE v1.1. [S01] Beam therefore has to prove not only that it is efficient relative to very large MoEs, but also that its full system economics hold up against aggressively sparse competitors.

Nemotron 3 Ultra and Inkling

NVIDIA’s Nemotron 3 Ultra has 550B total and 55B active parameters with up to 1M context, targeting long-running agents. [S34][S42] Thinking Machines’ Inkling has 975B total and 41B active parameters, 1M context, multimodal inputs, and already-released weights. [S35][S41]

Beam’s 23B active path is notably smaller than either, which could matter for throughput and cost. Real task-completion economics are still unknown.

Could Beam run locally on a Mac?

Not today in a meaningful way, because there is no public checkpoint to run. But the storage math gives a useful preview of what future quantized deployment might look like.

For 501 billion parameters, raw weight storage is approximately:

PrecisionFormulaRaw weights onlyPractical reading
BF16 / FP16501B × 2 bytes~1,002 GBfar beyond one Mac Studio
8-bit501B × 1 byte~501 GBleaves effectively no room for runtime overhead
6-bit501B × 0.75 bytes~375.75 GBcapacity may be possible on a 512GB machine, with limited headroom
4-bit501B × 0.5 bytes~250.5 GBcapacity is plausible on a 512GB machine
3-bit501B × 0.375 bytes~187.88 GBsmaller, but quality and runtime behavior are unknown

Author calculation. These numbers exclude quantization scales and metadata, KV cache, activations, runtime buffers, the operating system, and any MoE-specific overhead.

Apple’s current M5 Ultra Mac Studio can be configured with up to 512GB of unified memory and 1.2TB/s memory bandwidth. [S36][S37] That puts a future 4-bit Beam checkpoint inside the raw capacity envelope of one high-end Mac for the first time.

Capacity is not performance. We still do not know:

  • whether Reflection or partners will publish 4-bit, FP8, FP4, GGUF, or MLX-ready formats;
  • whether MLX, llama.cpp, or another Apple-silicon runtime will support Beam’s expert routing efficiently;
  • how much additional memory long contexts consume;
  • what prompt-prefill and decode speeds will look like;
  • whether 4-bit or 3-bit quantization preserves the model’s agentic quality;
  • whether partial expert offloading will be required.

Why “23B active” does not mean “runs like a 23B model”

A sparse MoE activates only a subset of experts for each token, but the model still needs access to the broader expert pool. That makes the compute path small while the storage footprint remains large.

A 128GB or 256GB Mac would therefore not be a straightforward target for a full 4-bit 501B checkpoint. A 512GB Mac Studio is much more interesting, but any concrete tokens-per-second claim before the checkpoint ships would be speculation.

Memory math for Beam on a Mac Studio with unified memory.

What can developers use today?

Reflection says an early Beam version is available to selected users through a waitlist. [S01]

Several technical publications that inspected Reflection’s beta documentation report an OpenAI-compatible Chat Completions interface and the model identifier Beam-501B-A23B, plus multiple reasoning-effort levels. [S11][S12][S15] Because those beta docs were not available as a stable public primary page during this review, production integrations should re-check them inside Reflection’s current developer console.

The following are still missing or not final enough for a production recommendation:

  • a public per-token price card,
  • downloadable weights,
  • final self-hosting hardware guidance,
  • final quantization formats,
  • the full model card,
  • the full technical report,
  • published safety-evaluation details,
  • broad independent benchmark reproduction.
Beam timeline: Early Access on October 5 and the planned open-weight release later in October 2026.

Who should watch Beam closely?

Beam matters most to:

  1. Coding-agent teams evaluating future open backends.
  2. Enterprises that need self-hosting and data control, where owning the model stack can matter more than the absolute top benchmark score.
  3. Researchers studying large-scale RL and MoE stability, because Reflection disclosed unusually large rollout and infrastructure numbers.
  4. Local-AI users with very high-memory workstations, because a 4-bit Beam could plausibly fit into a 512GB unified-memory class if the software stack supports it.

Should you wait for Beam?

For a production workload today, Beam is not the rational default over already-available GLM, Qwen, Kimi, DeepSeek, Nemotron, or Inkling deployments. [S29-S35]

Beam becomes much more compelling when three things happen:

  • Reflection actually publishes the promised Apache-2.0 weights;
  • independent coding and agent benchmarks broadly confirm the launch scores;
  • real serving data establishes latency, memory use, tokens per second, and cost per completed task.

For Mac users, add one more requirement: a working MLX or equivalent Apple-silicon implementation.

Bottom line: Reflection Beam is a credible preview, not a completed release

Beam is one of the more technically credible new open-weight announcements of 2026, not because 501B is a large number, but because Reflection pairs that capacity with a 23B active path, 23.8T-token pretraining, a 1M effective-context claim, and an RL campaign exceeding 100 million rollouts. [S01]

The correct status on October 6 is still preview, not completed release. The published benchmarks are promising but mostly vendor-reported. Local Mac feasibility is a capacity calculation, not a benchmark result. A 512GB M5 Ultra Mac Studio could theoretically hold a future 4-bit checkpoint with substantial headroom relative to the raw weights, but real-world feasibility depends on the missing checkpoint, quantization, runtime support, context memory, and measured throughput. [S01][S36]

The first tests to run after the weights ship

  • Verify checkpoint hashes, exact version, and final license.
  • Measure full model size in BF16/FP8/6-bit/4-bit formats.
  • Test peak unified-memory use on 256GB and 512GB Macs.
  • Separate prefill speed from decode speed.
  • Test 32K, 128K, and 256K contexts instead of quoting one best-case number.
  • Run the same agent harness against GLM-5.3, DeepSeek V4.1 Flash, and Qwen3.8-Max.
  • Compare cost or energy per completed task, not just token price or tokens per second.

Frequently Asked Questions

Can I download Reflection AI Beam now?

No. Reflection says the weights will be released later in October 2026. [S01]

Is Beam open source?

Reflection calls Beam an open-weight model and plans Apache 2.0 weights plus an open developer stack. Until the actual artifacts ship, “open weight” is the more precise term. [S01][S02]

Does Beam have a 1M-token context window?

Reflection says midtraining extended the model’s effective context to 1M tokens. The current beta serving limit is reported lower by multiple technical sources and is subject to change. [S01][S11][S12]

Is Beam better than Kimi K3 or DeepSeek V4.1 Flash?

Not universally. Reflection’s own table shows Kimi K3 and DeepSeek V4.1 Flash ahead on several coding and agent benchmarks. Beam’s strongest claim is efficiency, not across-the-board leadership. [S01]

Could Beam fit on a 512GB Mac Studio?

A raw 4-bit representation of 501B parameters is about 250.5GB, so capacity is plausible. That does not prove a usable implementation: runtime overhead, KV cache, expert routing, quantization quality, and software support still need measurement. [S36] --- **Source codes:** Full bibliographic details are in `data/source-register.csv`. The central primary source is Reflection AI’s October 5 launch post [S01].

Transparency

Sources and review basis

42

These primary and reference sources form the basis of the technical assessment. Vendor claims and external benchmarks are identified as such in the article.

  1. reflection.ai blog / introducing-beam
  2. reflection.ai reflection.ai
  3. reuters.com technology / nvidia-backed-reflection-unveils-first-ai-model-take-chinese-open-models-2026-10-05
  4. ft.com content / 353be303-3271-43f0-aefb-69b6ed7a0a6f
  5. axios.com 06 / reflection-mistral-open-weight-ai-models-china
  6. axios.com 04 / reflection-open-weight-ai
  7. techcrunch.com 05 / reflection-debuts-beam-a-open-weight-ai-model-to-rival-chinese-models-at-lower-compute-cost
  8. businessinsider.com what-is-beam-american-ai-model-deepseek-of-the-west-2026-10
  9. marketwatch.com story / this-new-ai-model-could-help-america-close-a-technological-gap-with-china-c1d74492
  10. theinformation.com briefings / reflection-ai-announces-first-open-source-model-beam
  11. datacamp.com blog / reflection-ai-beam
  12. aievals.app models / beam
  13. benchlm.ai models / reflection-beam
  14. s5labs.io insights / reflection-beam-501b-open-weights-announcement
  15. orcarouter.ai blog / reflection-beam-501b-explained
  16. cellcog.ai blog / reflection-beam-open-weight-model
  17. metirai.com blog / reflection-ai-beam-501b-open-weight-moe-model-benchmarks-2026
  18. thinkfacility.com blog / reflection-beam-open-weight-model-501b
  19. alphaxiv.org abs / 2610.introducing-beam
  20. ai-tldr.dev releases / reflection-beam
  21. ai-primer.com stories / reflection-beam-release
  22. swebench.com verified
  23. github.com zapier / AutomationBench
  24. github.com scaleapi / SWE-bench_Pro-os
  25. tbench.ai news / terminal-bench-2-1
  26. deepswe.datacurve.ai deepswe.datacurve.ai
  27. openai.com index / browsecomp
  28. longbench2.github.io longbench2.github.io
  29. alibabagroup.com en-US / document-2021044032125272064
  30. z.ai blog / glm-5.2
  31. z.ai blog / glm-5.3
  32. github.com MoonshotAI / Kimi-K3
  33. deepseek.com news / deepseek-v4-1-flash
  34. developer.nvidia.com blog / nvidia-nemotron-3-ultra-powers-faster-more-efficient-reasoning-for-long-running-agents
  35. thinkingmachines.ai news / introducing-inkling
  36. apple.com mac-studio / specs
  37. apple.com 08 / apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra
  38. artificialanalysis.ai articles / announcing-aa-lcr
  39. github.com allenai / IFBench
  40. api-docs.deepseek.com updates
  41. thinkingmachines.ai model-card / inkling
  42. build.nvidia.com nemotron-3-ultra-550b-a55b / modelcard