Urgent.News

What's breaking now, across thousands of outlets.

AI

One Self-Improving Model, Eleven Number-One Titles: What That Takes

TL;DR One self-improving model family now holds eleven public number-one benchmark records at the same time, across math, science, law, structured output, and decisions. The interesting part is not any single score, it is that a recursive self-improvement loop, bound to external verification, can push a whole spread of benchmarks to the top at once. The eleven, at a glance Math: AIME 2026 100%,…

A self-improving model, dubbed the S1MB number one model, has broken records in eleven distinct categories simultaneously, spanning mathematics, science, law, structured output, and decision-making. This feat is not attributed to a single model, but rather to a recursive self-improving (RSI) loop that pushes multiple benchmarks to the top at once.

The eleven records cover areas such as AIME 2026, HMMT 2026, GPQA Diamond, MMLU-Pro, MMMU-Pro, LEXam, LEXam-hard, ExtractBench, IFStruct, MDPBench, and S1MB Borda. A closer look reveals a 100% success rate in math and HMMT 2026, a 94.44% success in GPQA Diamond, and a 68.94% success in LEXam.

The RSI model operates by attempting problems, retaining those that pass an external verification, and training on these solutions. This method prevents the model from falling into its own errors, as it is tied to external checks like code execution or answer keys, rather than self-grading. The zero-token decision method, where the model reads a problem once and makes a calibrated decision without generating any text, further contributes to accurate and efficient decision-making.

Open-source enthusiasts can access the S1MB model, the ZTC method, and on-device versions through GitHub repositories. The decision model and method are available under the Apache-2.0 license, encouraging widespread use and collaboration.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

S1MB 1위: 0토큰으로 판정해 디시전 엔진 리더보드를 제패하다

요약 System One Mosaic Benchmark(S1MB) 디시전 엔진 리더보드에서 우리 모델 Darwin-27B-ZTC-v2가 102개 모델 중 1위입니다. Borda 점수 89.58, 태스크 평균 66.46. 핵심은 판정 방식입니다. 한 번의 순전파로, 생성 토큰 0개로 결정합니다. S1MB가 측정하는 것 S1MB는 타입드 결정 품질을 봅니다.

  • Darwin-27B-ZTC-v2 named top decision-making engine on S1MB leaderboard
  • Zero-Token Confidence method uses single forward pass, zero token generation
  • Model achieves 89.58 Borda score, 66.46 task average in S1MB benchmark

S1MB Número Uno: un juez de cero tokens que lideró el ranking de motores de decisión

Resumen En el ranking de motor de decisión de la System One Mosaic Benchmark (S1MB), nuestro modelo Darwin-27B-ZTC-v2 es número uno entre 102 modelos, con una puntuación Borda de 89.58 y un promedio…

  • Darwin-27B-ZTC-v2 named top decision-making model in S1MB
  • Zero-token generation process sets it apart from competitors
  • Achieved Borda score of 89.58 and average task score of 66.46

Engram Corruption: What Happens When a Skill Container Doesn't Own Its Payload

The setup In a modular AI framework like LivinGrimoire, behavior comes from small, swappable units called skills. A Brain holds them in lobes, and a skill's input() runs on every think cycle.

  • Skills are managed by higher-level AH skills, which handle skill management.
  • Engram snapshot process inadvertently includes payload skills, causing duplication issues.

More from Sunday 11 October →