Urgent.News

What's breaking now, across thousands of outlets.

AI

Eleven Number-One Records: Measuring a Model That Swept Math, Science, Law and Decisions

TL;DR A single self-improving model family holds eleven public number-one benchmark records at once, across mathematics, science, law, structured output, and decisions. This is a measurement note: the eleven, their scores, and why the spread matters more than any single win. The eleven records # Benchmark Field Result 1 AIME 2026 Math 100% (perfect) 2 HMMT 2026 Math 100% (perfect) 3 GPQA Diamond…

Eleven number-one records across math, science, law, structured output, and decisions have been achieved by a single self-improving model family - a significant milestone in the field of artificial intelligence. This includes perfect scores on the AIME 2026 math competition and HMMT 2026 math competition, as well as a 94.44% score on GPQA Diamond Science benchmark. The family's performance on structured output tasks, such as ExtractBench and IFStruct, are also noteworthy at 90.29% and 98.95% respectively.

The key here is that achieving top marks across these diverse fields is not due to tuning to a specific test, but rather indicative of a general method capable of handling different capabilities. This is particularly true when considering decision-oriented tasks like LEXam Law and MDPBench Decision process, where the scores are 68.94% and 83.65% respectively. The spread of these scores signals that the model's performance is not limited to a single area of expertise.

The methodology behind this achievement involves a recursive loop tied to external verification, ensuring the model's output is deterministic and reproducible. This is especially crucial in decision-making tasks where the output is typed. It is noteworthy that the model's performance on decision-making tasks has been evaluated using a zero-token method, which keeps the measurement honest and replicable.

For those interested in exploring this model further, the S1MB number one model is available on GitHub, along with the ZTC decision method. The underlying models can also be found on Hugging Face's repository. This groundbreaking achievement signifies a significant step forward in creating AI systems that can effectively handle a wide range of tasks, moving the field closer to the goal of general artificial intelligence.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

S1MB Number One: A Zero-Token Judge That Won the Decision-Engine Leaderboard

TL;DR On the System One Mosaic Benchmark (S1MB) decision-engine leaderboard, our model Darwin-27B-ZTC-v2 ranks number one among 102 models, with a Borda score of 89.58 and a task average of 66.46.

  • Darwin-27B-ZTC-v2 wins S1MB leaderboard with 89.58 Borda score
  • Zero-Token Confidence (ZTC) method uses single forward pass, no token generation
  • ZTC eliminates token drift, ensures consistent outputs for identical inputs

S1MB 1위: 0토큰으로 판정해 디시전 엔진 리더보드를 제패하다

요약 System One Mosaic Benchmark(S1MB) 디시전 엔진 리더보드에서 우리 모델 Darwin-27B-ZTC-v2가 102개 모델 중 1위입니다. Borda 점수 89.58, 태스크 평균 66.46. 핵심은 판정 방식입니다. 한 번의 순전파로, 생성 토큰 0개로 결정합니다. S1MB가 측정하는 것 S1MB는 타입드 결정 품질을 봅니다.

  • Darwin-27B-ZTC-v2 named top decision-making engine on S1MB leaderboard
  • Zero-Token Confidence method uses single forward pass, zero token generation
  • Model achieves 89.58 Borda score, 66.46 task average in S1MB benchmark

One Self-Improving Model, Eleven Number-One Titles: What That Takes

TL;DR One self-improving model family now holds eleven public number-one benchmark records at the same time, across math, science, law, structured output, and decisions.

  • S1MB model breaks records in eleven categories simultaneously
  • Recursive self-improving loop pushes multiple benchmarks to top
  • Zero-token decision method enhances accurate and efficient decision-making

More from Sunday 11 October →