Urgent.News

What's breaking now, across thousands of outlets.

AI

S1MB 1위: 0토큰으로 판정해 디시전 엔진 리더보드를 제패하다

요약 System One Mosaic Benchmark(S1MB) 디시전 엔진 리더보드에서 우리 모델 Darwin-27B-ZTC-v2가 102개 모델 중 1위입니다. Borda 점수 89.58, 태스크 평균 66.46. 핵심은 판정 방식입니다. 한 번의 순전파로, 생성 토큰 0개로 결정합니다. S1MB가 측정하는 것 S1MB는 타입드 결정 품질을 봅니다. 조건과 후보 선택지가 주어지면 모델이 맞는 것을 고르고 점수를 매깁니다. 텍스트 생성 벤치가 아니라 결정 벤치입니다. 출력이 타입드 정답키와 대조되므로 순위가 깔끔하고 재현됩니다. 애매한 문자열 매칭도, 판정용 2차 모델도 끼지 않습니다. 왜 0토큰이 유리한가 보통은 결정을 얻으려고 전체 생성 루프를 돌린 뒤 텍스트를 다시 파싱합니다. 느리고…

System One Mosaic Benchmark (S1MB) has crowned Darwin-27B-ZTC-v2 as the top decision-making engine among 102 models. The Borda score achieved is an impressive 89.58, with a task average of 66.46. The key to this victory lies in the S1MB's unique judging method. Instead of generating token outputs and parsing the text multiple times, it employs Zero-Token Confidence (ZTC), a method that only requires a single forward pass and zero token generation.

In this system, the model receives a text input, processes it through the final hidden state, and applies a corrected probe to instantly determine the correct answer. This approach ensures consistent results across different inputs and provides a straightforward ranking system, as the output directly matches the correct answer key. This method eliminates ambiguity and eliminates the need for secondary judgment models or repeated inference cycles.

The Darwin-27B-ZTC-v2 model has secured the top spot in the S1MB leaderboard, boasting an exceptional Borda score and a solid task average. For those interested in exploring the model and the ZTC method, the GitHub repositories for the S1MB benchmark, ZTC implementation, and the Darwin-27B-ZTC-v2 model weights are available. The open-source nature of the model under the Apache-2.0 license allows for widespread access and potential advancements in decision-making AI.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

S1MB Number One: A Zero-Token Judge That Won the Decision-Engine Leaderboard

TL;DR On the System One Mosaic Benchmark (S1MB) decision-engine leaderboard, our model Darwin-27B-ZTC-v2 ranks number one among 102 models, with a Borda score of 89.58 and a task average of 66.46.

  • Darwin-27B-ZTC-v2 wins S1MB leaderboard with 89.58 Borda score
  • Zero-Token Confidence (ZTC) method uses single forward pass, no token generation
  • ZTC eliminates token drift, ensures consistent outputs for identical inputs

One Self-Improving Model, Eleven Number-One Titles: What That Takes

TL;DR One self-improving model family now holds eleven public number-one benchmark records at the same time, across math, science, law, structured output, and decisions.

  • S1MB model breaks records in eleven categories simultaneously
  • Recursive self-improving loop pushes multiple benchmarks to top
  • Zero-token decision method enhances accurate and efficient decision-making

S1MB Número Uno: un juez de cero tokens que lideró el ranking de motores de decisión

Resumen En el ranking de motor de decisión de la System One Mosaic Benchmark (S1MB), nuestro modelo Darwin-27B-ZTC-v2 es número uno entre 102 modelos, con una puntuación Borda de 89.58 y un promedio…

  • Darwin-27B-ZTC-v2 named top decision-making model in S1MB
  • Zero-token generation process sets it apart from competitors
  • Achieved Borda score of 89.58 and average task score of 66.46

More from Sunday 11 October →