501B Stored, 23B Awake: Reflection's Beam and the Efficiency Turn in the Open-Weight Race
One-line: Reflection AI — a two-year-old Brooklyn startup with ~$4.7B raised and zero public models — just unveiled Beam : a 501B-parameter text-only MoE that wakes 23B per token , claims parity with Z.ai's GLM-5.2 at 3–4x less inference compute, and ships Apache 2.0 weights this month. On October 5, Reflection AI put its first public model on the table. The Brooklyn startup — founded 2024 by two…
Reflection AI, a two-year-old Brooklyn startup, recently unveiled Beam, a 501 billion-parameter text-only Mixture-of-Experts (MoE) model that activates 23 billion parameters per token. The startup claims Beam's per-token compute is 3-4 times less expensive than Z.ai's GLM-5.2, while claiming parity with it on reasoning benchmarks.
The company, founded in 2024 by former Google DeepMind researchers and backed by Nvidia, Sequoia, and Lightspeed, announced Beam and its weights on October 5. The startup's pretraining involved 23.8 trillion tokens, with a 1-million-token context window. The company asserts that Beam's architecture and serving stack contribute to its efficiency, but independent verification of the performance claims is lacking.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.