Urgent.News

What's breaking now, across thousands of outlets.

AI

S1MB Number One: A Zero-Token Judge That Won the Decision-Engine Leaderboard

TL;DR On the System One Mosaic Benchmark (S1MB) decision-engine leaderboard, our model Darwin-27B-ZTC-v2 ranks number one among 102 models, with a Borda score of 89.58 and a task average of 66.46. What makes the result different is how it decides: a single forward pass, zero generated tokens. What S1MB tests S1MB measures typed-decision quality. Each item gives a condition and candidate choices,…

The System One Mosaic Benchmark (S1MB) has crowned Darwin-27B-ZTC-v2 as the top model, reaching number one out of 102 contenders. Darwin-27B-ZTC-v2 scored an impressive Borda score of 89.58 and a task average of 66.46 across all tests. What makes this achievement unique is the model's decision-making process: it relies on a single forward pass without generating any tokens.

This "zero-token decision" method is known as ZTC (Zero-Token Confidence), which sets it apart from other systems that typically generate text before extracting a decision. By bypassing the generation loop, ZTC eliminates the slow, non-deterministic nature of traditional approaches and prevents token drift. The ZTC method works by reading the input in one forward pass, extracting the final-layer hidden state, and applying a calibrated probe to generate the decision directly.

This approach guarantees the same output for identical inputs, a crucial factor for both measurement and production environments. The complete S1MB number one model, ZTC method, and inference code are all available on GitHub. Additionally, the weights for the Darwin-27B-ZTC-v2 model can be accessed via Hugging Face.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Nvidia discusses boosting Reflection investment, computing deal

Reflection AI debuted in October its first open-weight model, a market dominated by Chinese startups.

  • Nvidia in early-stage talks to increase investment or computing resources in Reflection AI.
  • Potential deal structures include acqui-hire, additional chips, or equity stake increase.
  • Reflection launched open-weight model Beam in October, targeting business clients and developers.

S1MB 1위: 0토큰으로 판정해 디시전 엔진 리더보드를 제패하다

요약 System One Mosaic Benchmark(S1MB) 디시전 엔진 리더보드에서 우리 모델 Darwin-27B-ZTC-v2가 102개 모델 중 1위입니다. Borda 점수 89.58, 태스크 평균 66.46. 핵심은 판정 방식입니다. 한 번의 순전파로, 생성 토큰 0개로 결정합니다. S1MB가 측정하는 것 S1MB는 타입드 결정 품질을 봅니다.

  • Darwin-27B-ZTC-v2 named top decision-making engine on S1MB leaderboard
  • Zero-Token Confidence method uses single forward pass, zero token generation
  • Model achieves 89.58 Borda score, 66.46 task average in S1MB benchmark

More from Sunday 11 October →