Accéder au contenu principal

Beam: Reflection AI's 501B Open-Weight Coding Model

Reflection AI's Beam matches GLM 5.2 on reasoning with 3-4x less compute, but trails Kimi K3 on agentic coding. Here's how Reflection AI's open-weight model performs.
6 oct. 2026  · 12 min lire

Explorez l'IA

ChatGPTClaudePerplexity

Open-weight model releases have been dominated for a while by Chinese labs, with the GLM, Kimi, Qwen, and DeepSeek families taking the top scores on agentic coding evaluations.

Reflection AI's Beam, announced on October 5, 2026, comes with an explicit pitch: advance what the lab calls the "Western open-weight frontier", and compete on inference efficiency rather than raw capability.

In this article, I'll cover what Reflection AI's Beam is, the features that matter for day-to-day work, how it compares to Kimi K3, DeepSeek V4.1 Flash, and the GLM 5.x line on the benchmarks, and what you can and cannot do with it today. I also recommend checking out our guide to another new model, Kolibri 1, and our guides to Kimi K3 and DeepSeek V4.1 Flash vs Gemini 3.8 Flash.

TL;DR

  • Beam is Reflection AI's first open-weight model, a 501B-parameter mixture-of-experts (MoE) model that uses only 23B parameters for each token it generates.
  • The selling point is intelligence per token, not peak capability: Reflection says it matches GLM 5.2 on reasoning while using 3 to 4 times less inference compute.
  • Kimi K3 still leads on raw capability, and DeepSeek V4.1 Flash posts the top scores on individual benchmarks like Terminal-Bench v2.1 and DeepSWE.
  • Weights are promised under Apache 2.0 later in October 2026, so you cannot self-host, quantize, or fine-tune it yet.

What Is Reflection AI's Beam?

Beam is a 501-billion-parameter sparse mixture-of-experts model from Reflection AI, built for coding, reasoning, and agentic workloads. Only 23 billion of those parameters are active on any given token.

It is the lab's first open-weight release and the first model in a planned series.

Under the hood, Beam has 52 layers, splits its knowledge across many small specialist "experts" (only a few run per token), and mixes attention that looks at nearby text with attention that looks across the whole input.

The sparsity is the whole argument.

A 2T-plus class model like Qwen 3.8-Max activates far more parameters per token, so every generated token costs more compute.

Reflection's bet is that a 23B active footprint, pushed hard with reinforcement learning (RL), gets close enough on coding and agentic work to make the cost difference worth it.

Training scale is where Reflection spent its money. Pretraining on 23.8 trillion tokens took under 4 weeks on 6,144 NVIDIA GB300 GPUs, with the cluster doing useful work 92.3% of the time toward the end of the run and 9 restarts from earlier checkpoints along the way.

The RL phase then ran separately on 10,500 GB300s for 4 weeks. It produced more than 100 million rollouts (complete attempts at a task) across roughly 1 million curated coding, agentic, and STEM tasks, using around 1.3 billion sandboxed environments to run and grade them.

Reflection also reports that capability kept improving as RL compute increased, with no plateau by the end of the run. That's the claim I would most like to see tested externally, because it underpins the lab's statement that it is already training the next model in the series.

Reflection AI Beam Model Card

Beam Key Features

Beam is text-only, so there is no image or audio input.

What it does offer is a set of controls and capabilities aimed squarely at people running coding agents.

A reasoning effort dial you set per request

Beam exposes a reasoning effort parameter, so you choose how many thinking tokens a task gets. The API accepts low, medium (the default), high, xhigh, and max, but reasoning cannot be switched off entirely.

Lower settings favor short responses; higher settings allow longer reasoning on harder problems.

The dial is a direct result of how Beam was trained.

During RL, Reflection used an adjustable length penalty that rewarded correct solutions while discouraging extra tokens. It reports that early in training, performance improved while responses actually got shorter.

Later, as agentic ability grew, responses got longer again, and those extra tokens paid for themselves.

1M-token context for whole-repository work

Midtraining extends Beam's effective context to 1 million tokens, trained on structured code repositories, long-horizon tasks, and long-form documents.

There are two caveats. RL rollouts ran with a maximum context of 256K tokens, so Beam's agentic behavior was learned at a quarter of that length, and the beta API currently caps requests at 262,144 tokens (input plus output), with output limited to 131,072 tokens.

Long context on paper and long context in practice are different things. Beam scores 65.5 on LongBench v2 against Qwen 3.8-Max's 66.3 and GLM 5.2's 64.0, and 79.3 on AA-LCR, where Kimi K3 reaches 88.7. That makes Beam useful for long inputs, but not class-leading.

Agentic skills that carry over to tasks it was never trained on

The most interesting result in the announcement has nothing to do with a leaderboard.

During a training phase covering reasoning, software engineering, and terminal tasks, with no browsing tasks in the RL mix at all, Beam's browsing scores improved anyway.

Given web access, the model worked out on its own that it could query other language models and call optical character recognition (OCR) APIs to read documents it could not see.

Reflection's demos push this further: a live NYC subway map built from public MTA data, a 3D p5.js free-fall game built entirely through text, and a Gemma 4 fine-tuning notebook assembled inside OpenCode from Unsloth documentation, which lifted Gemma's held-out test accuracy by 66.5%.

One generalization test from the announcement is worth singling out because it's hard to have seen in training data.

Reflection recreated a viral grid puzzle posted to X a few days before launch, asking Beam to classify each of 16,200 points on a fixed 180x90 world grid as land or water. Beam got 95.5% of them right, between Opus 5 at 92.5% and Fable 5 at 97.8%.

Asynchronous RL with up to one day of lag

Beam was trained with asynchronous RL, meaning the model keeps generating rollouts while training updates happen in parallel, rather than waiting for each other. A single long rollout can therefore contain tokens produced by several different versions of the model.

Reflection reports stable learning even when a batch's oldest sample came from a model version 107 updates behind the current one, roughly one day out of date.

The infrastructure numbers are worth a look if you build RL systems:

  • 110K concurrent rollouts on average, with up to 170K concurrent sandboxes
  • About 12 seconds (median) for new weights to reach the inference servers
  • 71 inference incidents handled without stopping the training job

Open weights under Apache 2.0, plus the stack around them

Reflection says it will publish Beam's weights under Apache 2.0 later in October 2026, along with the technical report, model card, documentation, and tooling for running, evaluating, and fine-tuning the model.

It also plans to open-source the internal safety evaluations it built, which is rarer than the weight drop itself.

Apache 2.0 is about as permissive as it gets: commercial use, redistribution, and modification, with no user or revenue thresholds.

That matters because Kimi K3 and DeepSeek V4.1 Flash have set expectations for what an open model should let you do, and a restrictive license would have taken Beam out of the conversation.

How Does Beam Perform on the Benchmarks?

Beam lands roughly where GLM 5.2 sits on most evaluations, well ahead of Inkling and Nemotron 3 Ultra, and clearly behind Kimi K3, GLM 5.3, and DeepSeek V4.1 Flash on the hardest agentic coding tasks.

Reflection is upfront about this, framing Beam's advantage as efficiency rather than peak capability.

All figures below are vendor-reported, with rivals' scores sourced by Reflection from Artificial Analysis and DataCurve.

Reflection beam-benchmarks-efficiency

Coding and agentic workflows

Beam's headline coding number is 80.9 on SWE-bench Verified, against 77.6 for Inkling and 70.7 for Nemotron 3 Ultra.

SWE-bench Verified is a human-checked set of real GitHub issues where a model has to write a fix that passes the project's own tests, and it remains the first benchmark most practitioners check.

Elsewhere, the picture is less flattering.

Against the models you would actually be choosing between, Beam trails on every agentic coding benchmark where a comparison score exists.

Benchmark Beam GLM 5.2 GLM 5.3 Kimi K3 DeepSeek V4.1 Flash
Terminal-Bench v2.1 80.1 81.0 88.2 88.3 90.6
DeepSWE v1.1 44.4 44.0 61.0 68.0 74.2
SWE-bench Pro v2-Hard 77.2 NR 84.3 88.2 NR
SWE Atlas Codebase QnA 34.6 NR 61.0 68.0 NR

NR = not reported.

The DeepSWE and SWE Atlas gaps are large.

DeepSWE v1.1 tests multi-step agentic software engineering rather than single fixes, and a 29.8-point deficit to DeepSeek V4.1 Flash is not a rounding error.

On SWE-bench Pro v1, Beam's 65.5 does beat GLM 5.2's 62.1 and sits just under Qwen 3.8-Max's 67.7.

Reasoning and knowledge tasks

On Humanity's Last Exam (HLE) without tools, Beam scores 36.2 against 40.5 for GLM 5.2, 42.3 for GLM 5.3, 43.6 for Qwen 3.8-Max, and 46.9 for Kimi K3.

HLE is a set of expert-written questions designed so that search and memorization do not help much, and it tends to reward the bigger models in this group.

The rest of the reasoning suite is tighter:

  • AIME 2026: Beam 97.8, Inkling 97.1, GLM 5.2 99.2
  • GPQA Diamond: Beam 90.5, GLM 5.3 91.7, Kimi K3 93.5
  • SciCode: Beam 49.7, Qwen 3.8-Max 52.1, Kimi K3 58.7
  • CritPt: Beam 15.6, GLM 5.3 19.1, Kimi K3 23.4

Almost every model here scores near the top on AIME, so it no longer separates them, and a 3-point GPQA Diamond gap to Kimi K3 is not what decides a production deployment.

CritPt, a benchmark of research-level physics problems, is where the difference bites, with Kimi K3 scoring around 50% higher than Beam.

Tool calling and search

Beam's tool-calling results are its most uneven set.

It scores 37.0 on AutomationBench (public set), ahead of GLM 5.2 at 26.2 but behind Kimi K3 at 46.7 and DeepSeek V4.1 Flash at 54.8. On MCP Atlas, it scores 78.7 against 77.8 for GLM 5.2 and 84.5 for Qwen 3.8-Max.

On τ³-bench banking, Beam reaches 38.0 against 37.1 for both GLM 5.2 and Kimi K3, with Qwen 3.8-Max well clear at 55.2. Search is the weakest area: 77.4 on BrowseComp and 80.1 on DeepSearchQA with context management, against Kimi K3's 91.2 and 95.0. If your agent lives on the open web, that gap is the one to care about.

Inference efficiency

Inference efficiency is the number Reflection actually wants you to look at.

On Reflection's HLE efficiency chart, Beam reaches 35.3% at an estimated 0.82 PFLOP (quadrillion floating-point operations) of compute per attempt, while Inkling needs around 4.1 PFLOP to reach roughly 26% and Nemotron 3 Ultra around 1.9 PFLOP for about 21%.

Reflection's broader claim is that Beam matches GLM 5.2 on advanced reasoning while using 3 to 4 times less inference compute, with a bigger gap against the 2T-plus Qwen 3.8-Max family.

Read the method carefully: compute is estimated as 2 times active parameters times the average number of generated tokens per attempt. It leaves out the cost of processing the prompt, the extra attention cost of long contexts, and serving overhead.

It's an approximate compute comparison, not a measured deployment cost, and because it only counts active parameters, it naturally favors sparse models like Beam.

How Was Beam Aligned?

Reflection trained a separate safety and alignment model from the same pretrained checkpoint, then combined it with the main RL-trained model using multi-teacher on-policy distillation (MOPD), where a single student model learns from both "teacher" models at once.

The alignment principles are organized into three tiers:

  • Hard rules the model should not break
  • Qualities it should consistently show, such as acknowledging uncertainty
  • A default interaction style described as direct, thorough, and proactive

Safety training used deliberative alignment, which teaches the model to reason through its safety policy before answering. An adversarial loop generated attacks against each round's checkpoint and fed the successful ones back into the next round of supervised fine-tuning data.

Reflection also reports it could predict how much RL would improve hard-to-grade behaviors with a correlation of r = 0.79, compared with 0.46 using a simpler Best-of-N baseline. Safety evaluation results are promised in the technical report rather than published now.

Beam Pricing and Availability

There is no published price for Beam. Reflection's developer documentation had no per-token rate card as of October 6, 2026, which makes any cost comparison against hosted Kimi K3 or DeepSeek V4.1 Flash endpoints impossible right now.

Access is through an API beta opened gradually via a waitlist.

The documented base URL is https://api.reflection.ai/openai/v1, and it is OpenAI-compatible in a narrow sense: Chat Completions and model listing are supported, while Responses, Embeddings, Files, and Batch are not. Beta limits and behavior are explicitly subject to change.

Distribution is thin so far. When I checked on October 6, 2026, Beam was not on Ollama, not in the OpenRouter catalog, not in the models.dev catalog used by OpenCode, and no weights were listed on Hugging Face.

Reflection says it will launch with distribution partners and open-source harness integrations alongside the weight release.

How to Access Reflection AI's Beam

Beam's weights are not out yet.

Reflection has committed to an Apache 2.0 release later in October 2026, covering the weights, technical report, model card, and tooling for running, evaluating, and fine-tuning the model. Until that lands, you cannot self-host it, quantize it, measure how much GPU memory it really needs, or reproduce any published score.

What you can do today is join the beta waitlist and call the hosted API.

The model ID in the developer documentation is Beam-501B-A23B, and because the endpoint speaks Chat Completions, the official OpenAI Python client works once you change the base URL.

from openai import OpenAI

client = OpenAI(
    base_url="https://api.reflection.ai/openai/v1",
    api_key="YOUR_REFLECTION_KEY",
)

response = client.chat.completions.create(
    model="Beam-501B-A23B",
    reasoning_effort="medium",  # low, medium, high, xhigh, or max
    messages=[{"role": "user", "content": "Port this module to the v2 payments API."}],
)
print(response.choices[0].message.content)

If you want to practice this pattern with a model you can access right now, our GPT-6 Sol API tutorial uses the same client code, including tool permissions and cost tracking for an agentic migration task.

Final Thoughts

Beam is a credible open-weight entry, and the efficiency framing is the honest one: Kimi K3 and DeepSeek V4.1 Flash beat it on the agentic coding benchmarks that matter most, so the case for Beam has to be cost per solved task rather than capability.

I would not plan around it yet.

Without published weights, a price, or a single independent reproduction, every claim here comes from the lab that trained the model, and the compute comparison leaves out prompt processing and serving costs. Check back when the Apache 2.0 weights land and Artificial Analysis publishes its own numbers.

If you want to get hands-on with the open-weight stack in the meantime, our AI Fundamentals skill track is a good place to build the groundwork.

FAQs

What is Reflection AI's Beam model?

Beam is Reflection AI's first open-weight model, announced on October 5, 2026. It is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token, pretrained on 23.8 trillion tokens and tuned with over 100 million reinforcement learning rollouts. It targets coding, reasoning, and agentic workloads, and is text-only.

How does Beam compare to Kimi K3 and DeepSeek V4.1 Flash?

Beam trails both on agentic coding. On Terminal-Bench v2.1 it scores 80.1 against Kimi K3's 88.3 and DeepSeek V4.1 Flash's 90.6, and on DeepSWE v1.1 it scores 44.4 against 68.0 and 74.2 respectively. Reflection's argument is efficiency rather than peak capability: Beam claims GLM 5.2-level reasoning at 3 to 4 times less inference compute.

Where can I download Beam's weights?

Nowhere yet. Reflection AI says it will publish the weights under an Apache 2.0 license later in October 2026, together with the technical report, model card, and tooling. As of October 6, 2026, no Beam weights were listed on Hugging Face, and Beam was not in the OpenRouter catalog.

How much does Beam cost to use?

No per-token pricing has been published. Reflection AI's developer documentation carried no rate card and no model-listing prices as of October 6, 2026, so hosted cost comparisons against other open models are not possible yet. Access is currently through a free-to-join API beta waitlist.

What is Beam's API model ID and endpoint?

The model ID in Reflection's developer documentation is Beam-501B-A23B, and the base URL is https://api.reflection.ai/openai/v1. The endpoint supports Chat Completions and model listing, so the OpenAI Python client works with a changed base URL. Completions, Embeddings, and Batches are not supported.

Matt Crabtree's photo
Author
Matt Crabtree
LinkedIn

A senior editor in the AI and edtech space. Committed to exploring data and AI trends.  

Sujets
Artificial Intelligence
Large Language Models

Top DataCamp Courses 

Cours

Coder avec l’aide de l’IA pour les développeurs

1 h 30 min
10.5K
Améliorez vos compétences en codage grâce à l'IA : guidez votre assistant de codage pour qu'il écrive, teste et documente efficacement le code.
Voir les détailsRight Arrow
Commencer Le Cours
Voir plusRight Arrow
Contenus associés

blog

Kimi K3: Moonshot AI's Newest and Best Open-Source Model

Read about Kimi K3 : Everything we know about Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-source LLM and the largest open-weight model released to date. See benchmarks, pricing, and API features.
Josef Waples's photo

Josef Waples

11 min

blog

GLM-5.2: Features, Setup, Benchmarks, and Model Switching Guide

Z.ai's GLM-5.2 ships with a 1M token context window, two reasoning effort levels, and free access across all GLM Coding Plan tiers.
Matt Crabtree's photo

Matt Crabtree

11 min

blog

Kolibri 1: Aleph Alpha's Sovereign Open-Weight LLM

Aleph Alpha has released Kolibri 1, a 78B mixture-of-experts model with 3.46B active parameters, Apache 2.0 weights, and a 1M-token context, trained from scratch in Germany.
Matt Crabtree's photo

Matt Crabtree

14 min

blog

GLM-5 vs GPT-5.3-Codex: Which AI Model Wins for Agent Workflows?

We compare GLM 5 vs GPT 5.3 Codex for AI agent workflows, analyzing architecture, benchmarks, deployment choices, and costs to guide your model selection.
Brian Mutea's photo

Brian Mutea

15 min

Tutoriel

Reflection Llama-3.1 70B: Testing & Summary of What We Know

Reflection Llama-3.1 70B, trained with Reflection-Tuning, claims to surpass GPT-4o and Claude 3.5 Sonnet but has faced reproducibility and verification issues so far.
Ryan Ong's photo

Ryan Ong

8 min

Tutoriel

Run GLM-5 Locally For Agentic Coding

Run GLM-5, the best open-weight AI model, on a single GPU with llama.cpp, and connect it to Aider to turn it into a powerful local coding agent.
Abid Ali Awan's photo

Abid Ali Awan

8 min

Voir PlusVoir Plus