Reflection unveils Beam, a 501B open-weight MoE model trained on 10.5K GB300s, with weights promised later this month
- Reflection introduced Beam, its first open-weight model: a sparse Mixture-of-Experts with 501B total parameters and 23B active, built for coding, reasoning, and agentic workloads.
- Pretraining ran on 23.8T tokens from the web plus proprietary licensed datasets, and the high-compute RL run produced over 100M rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks.
- The model is still in final red-teaming and evaluation; only an early-access signup is open, and the weights, technical report, model card, and developer artifacts are promised later this month.
- Beam scores 80.9 on SWE-Bench Verified, 80.1 on Terminal Bench v2.1, 97.8 on AIME 2026 and 90.5 on GPQA Diamond, and Reflection says it matches GLM 5.2 on advanced reasoning using 3 to 4 times less inference compute.
- Reflection states Kimi K3 stays ahead on raw capability and pitches Beam's edge as inference efficiency; the benchmark table includes GLM 5.3 and DeepSeek V4.1 Flash, but the charts leave both out.
Hacker News opinions
Early access signup, no weights, no tech details, and a 'proprietary licensed dataset' they won't name. Talk is cheap. Publish the weights and the HF repo or don't bother.
'Proprietary dataset' usually means either they can't show it, or there's nothing special in it and they want it to sound like a secret ingredient while there's none.
If I had a data center full of GPUs I'd just buy the data. scale.com/data-engine exists, FineWeb is 50TB on HuggingFace. Data sourcing isn't the mystery people make it out to be.
There's a lot of open research on pretraining and RL data mixtures, check Datalogy, Nvidia Nemotron, Ai2 and the recent Aleph Alpha model if you actually want to learn.
That world map generalization test they're bragging about is not a few days old. It showed up on LessWrong back in August 2025, so the 'could not appear in training data' claim is wrong.
The age of the puzzle barely matters anyway. Every model can search the web now.
Look at the performance chart: the better open models sit behind the fold so it looks like Beam wins. It doesn't. The only thing I took away was that I should try DeepSeek 4.1 Flash.
It's bigger than DeepSeek v4.1 Flash, more expensive to run, and worse on every measured metric. What am I missing here?
It's a new US entrant in this weight class, and that's the real selling point. 'Advances the Western open-weight frontier' is about government contracts that forbid foreign models, not about individuals.
At least it isn't distilled from every major American provider. More than some labs can say.
I want to know what hardware this runs on. Optimizing for inference speed sounds good for small boxes, but 501B total parameters could still be brutal.
Chinese open models are smaller and still ahead right now. The West is behind, but they showed up, and it's about relative pace at this point.
Credit where it's due, they admit Kimi K3 is ahead on raw capability instead of pretending otherwise. That's refreshing.
No AA Index, no Arena results, and GLM 5.3 and DeepSeek 4.1 Flash are in the table but not the charts. That omission tells you something.