Skip to content

Reflection AI Opens Beam Waitlist: 501B MoE With 23B Active, Vendor-Reported SWE-Bench Verified 80.9

Reflection AI announced Beam on 5 Oct 2026: a 501B sparse MoE with 23B active for coding and agents. Waitlist open; Apache 2.0 weights later this month. Benchmarks are Reflection-reported.

Official Reflection AI Introducing Beam blog OG image from reflection.company

Reflection AI announced Beam on 5 October 2026 as its first open-weight model. The company describes Beam as a sparse Mixture-of-Experts stack with 501 billion total parameters and 23 billion active per token, aimed at coding, reasoning, and agentic workloads.

Weights are not public yet. Reflection says Beam is still in final red-teaming and evaluations, early access is waitlist-only, and the weights, technical report, model card, and developer artifacts are planned later this month under Apache 2.0. There is no public API price on the launch page.

On 8 October 2026 Reflection updated the benchmark tables on the same post. The scores below are Reflection-reported from that Oct 8 update. They are not independent audits.

What Reflection Says Beam Is

Beam’s size story is efficiency at inference, not raw parameter count alone. With 23B active parameters per token, Reflection argues the model can sit near larger open models on coding and agent tasks while using less generation compute. The company says Beam advances the Western open-weight frontier and is competitive with larger open models such as GLM 5.2, while still trailing some Chinese open models like Kimi K3 on several coding benches where those models remain ahead on raw capability.

Training claims on the launch post are large and specific. Reflection says Beam was pretrained on 23.8 trillion curated tokens from the web and proprietary licensed datasets. Separately, it says a high-compute RL run produced more than 100 million rollouts on 10.5K NVIDIA GB300 GPUs over about four weeks, with training and grading using approximately 1.3 billion sandboxes. Reflection calls this one of the largest-scale RL runs by any open lab to date. That framing is the company’s.

Context and control knobs matter for builders. Midtraining extends effective context to 1M tokens, per Reflection. Users also get a reasoning effort parameter so lower settings favor shorter answers and higher settings allow longer reasoning for harder tasks. There is no USD list price yet because public serving is still waitlist and weight release.

Vendor Benchmarks From the Oct 8 Tables

Reflection published side-by-side tables against Inkling, Nemotron 3 Ultra, GLM 5.2 / 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash. NR means not reported. Treat every number as Reflection-reported.

Benchmark (Reflection-reported) Beam Notes from Reflection’s table
DeepSWE v1.1 44.4 Behind Kimi K3 (68.0) and DeepSeek V4.1 Flash (74.2)
SWE Bench Verified 80.9 Ahead of Inkling (77.6) and Nemotron 3 Ultra (70.7) in the same table
Terminal Bench v2.1 80.1 Near GLM 5.2 (81.0); behind Kimi K3 (88.3) and DeepSeek V4.1 Flash (90.6)
AIME 2026 97.8 Close to Inkling (97.1); GLM 5.2 listed at 99.2
GPQA Diamond 90.5 Near DeepSeek V4.1 Flash (90.9); Kimi K3 listed at 93.5

Reflection also claims Beam matches frontier-ish reasoning scores while using about 3 to 4 times less inference compute than GLM-5.2 on advanced reasoning, with even larger gaps versus 2T-plus parameter models such as Qwen 3.8-Max. Those efficiency charts use estimates from Artificial Analysis and DataCurve data and approximate FLOPs from active parameters times generated tokens. They are approximate comparisons, not measured serving bills.

Read the leaderboard the same way you read other vendor open-weight drops this year, from Aleph Alpha’s Kolibri to Cloudflare’s Clef pair. Useful as a map of what the lab wants you to notice. Not a substitute for your own harness.

RL Scale, Midtraining, and Safety Claims

Reflection puts high-compute RL at the center of Beam’s story. It says the campaign used about one million coding, agentic, and STEM environments, sustained averages around 110K concurrent rollouts, and kept learning stable even with day-old policy staleness through new asynchronous RL algorithms. Capability plots on Terminal-Bench 2.1, HLE, and DeepSWE are shown rising with cumulative rollouts, with no plateau claimed by the end of the run.

Pretraining ran end-to-end in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs, according to the post, with in-house scheduling, health monitoring, and silent data corruption handling. Reflection says late-run goodput reached 92.3%. Midtraining is framed as building a stronger prior for RL, including long-horizon tasks and the jump to 1M-token effective context.

Safety work used a separate SFT and RL teacher merged via multi-teacher on-policy distillation, plus deliberative alignment style training. Detailed safety eval results are promised in the upcoming technical report, with some internal safety evals to be open-sourced. Until that report ships, those claims stay vendor narrative.

What It Means for Indian Developers

If you already self-host open MoE stacks, Beam is interesting once weights land under Apache 2.0. The practical bet is coding and agent workloads where 23B active parameters may fit tighter inference budgets than denser frontier open models. Until the weight drop and model card arrive, the only action is the waitlist, not production planning against a price sheet.

Teams in India shipping SaaS copilots, terminal agents, or SWE benches should treat the Oct 8 scores as a shopping list for their own evals: DeepSWE, SWE-Bench Verified, Terminal Bench, plus domain suites in Indian English and local enterprise codebases. Compare against other recent open drops such as Thinking Machines’ Inkling and bilingual MoE options like Kolibri when sovereignty or on-prem constraints matter more than peak coding score.

USD pricing is still missing. Reflection has not published hosted token rates. Budget models should assume self-host GPU cost only after artifacts ship, and keep a closed API fallback until third-party runners publish real throughput and price.

Confirmed vs Unconfirmed

Item Status
Announcement date 5 Oct 2026; tables updated 8 Oct 2026 Confirmed on Reflection’s blog
501B total / 23B active sparse MoE; coding and agent focus Vendor claim on the launch post
23.8T pretrain tokens; 100M+ RL rollouts; 10.5K GB300; ~1.3B sandboxes Vendor claim; not independently verified here
DeepSWE 44.4, SWE-Bench Verified 80.9, Terminal Bench v2.1 80.1, AIME 2026 97.8, GPQA Diamond 90.5 Reflection-reported Oct 8 tables
Early access waitlist now; Apache 2.0 weights and artifacts later this month Stated plan on the launch post; not completed yet
Public API USD price Not disclosed
Independent third-party evals of Beam Not available in this article

Frequently Asked Questions

Is Beam downloadable today?

No. Reflection says you can join an early access waitlist now. Weights, the technical report, model card, and developer artifacts are planned later this month under Apache 2.0.

How big is Beam, and what is it for?

Reflection lists 501 billion total parameters with 23 billion active per token in a sparse MoE. The stated focus is coding, reasoning, and agentic workloads, with a controllable reasoning effort setting and midtraining to 1M-token context.

Should I trust the SWE-Bench and Terminal Bench scores?

Use them as Reflection-reported figures from the Oct 8 table update. They are useful for ranking what the lab claims against Kimi K3, GLM, Qwen, and DeepSeek in the same post. Run your own harness before any production decision.

How does Beam compare with Chinese open models?

Reflection itself says models like Kimi K3 remain ahead on several raw coding capability benches, and pitches Beam on inference efficiency instead. That is the company’s framing, not an external audit.

Is there a public price?

Not on the launch page. Until hosted pricing or partner SKUs appear, assume waitlist access and a future self-host path after the Apache 2.0 weight release.

Share this article

1 comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Loading the next article…

Continue reading