One Self-Improving Model, Eleven Number-One Titles: What That Takes
TL;DR One self-improving model family now holds eleven public number-one benchmark records at the same time, across math, science, law, structured output, and decisions. The interesting part is not any single score, it is that a recursive self-improvement loop, bound to external verification, can push a whole spread of benchmarks to the top at once. The eleven, at a glance Math: AIME 2026 100%,…
A self-improving model, dubbed the S1MB number one model, has broken records in eleven distinct categories simultaneously, spanning mathematics, science, law, structured output, and decision-making. This feat is not attributed to a single model, but rather to a recursive self-improving (RSI) loop that pushes multiple benchmarks to the top at once.
The eleven records cover areas such as AIME 2026, HMMT 2026, GPQA Diamond, MMLU-Pro, MMMU-Pro, LEXam, LEXam-hard, ExtractBench, IFStruct, MDPBench, and S1MB Borda. A closer look reveals a 100% success rate in math and HMMT 2026, a 94.44% success in GPQA Diamond, and a 68.94% success in LEXam.
The RSI model operates by attempting problems, retaining those that pass an external verification, and training on these solutions. This method prevents the model from falling into its own errors, as it is tied to external checks like code execution or answer keys, rather than self-grading. The zero-token decision method, where the model reads a problem once and makes a calibrated decision without generating any text, further contributes to accurate and efficient decision-making.
Open-source enthusiasts can access the S1MB model, the ZTC method, and on-device versions through GitHub repositories. The decision model and method are available under the Apache-2.0 license, encouraging widespread use and collaboration.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.