S1MB 1위: 0토큰으로 판정해 디시전 엔진 리더보드를 제패하다
요약 System One Mosaic Benchmark(S1MB) 디시전 엔진 리더보드에서 우리 모델 Darwin-27B-ZTC-v2가 102개 모델 중 1위입니다. Borda 점수 89.58, 태스크 평균 66.46. 핵심은 판정 방식입니다. 한 번의 순전파로, 생성 토큰 0개로 결정합니다. S1MB가 측정하는 것 S1MB는 타입드 결정 품질을 봅니다. 조건과 후보 선택지가 주어지면 모델이 맞는 것을 고르고 점수를 매깁니다. 텍스트 생성 벤치가 아니라 결정 벤치입니다. 출력이 타입드 정답키와 대조되므로 순위가 깔끔하고 재현됩니다. 애매한 문자열 매칭도, 판정용 2차 모델도 끼지 않습니다. 왜 0토큰이 유리한가 보통은 결정을 얻으려고 전체 생성 루프를 돌린 뒤 텍스트를 다시 파싱합니다. 느리고…
System One Mosaic Benchmark (S1MB) has crowned Darwin-27B-ZTC-v2 as the top decision-making engine among 102 models. The Borda score achieved is an impressive 89.58, with a task average of 66.46. The key to this victory lies in the S1MB's unique judging method. Instead of generating token outputs and parsing the text multiple times, it employs Zero-Token Confidence (ZTC), a method that only requires a single forward pass and zero token generation.
In this system, the model receives a text input, processes it through the final hidden state, and applies a corrected probe to instantly determine the correct answer. This approach ensures consistent results across different inputs and provides a straightforward ranking system, as the output directly matches the correct answer key. This method eliminates ambiguity and eliminates the need for secondary judgment models or repeated inference cycles.
The Darwin-27B-ZTC-v2 model has secured the top spot in the S1MB leaderboard, boasting an exceptional Borda score and a solid task average. For those interested in exploring the model and the ZTC method, the GitHub repositories for the S1MB benchmark, ZTC implementation, and the Darwin-27B-ZTC-v2 model weights are available. The open-source nature of the model under the Apache-2.0 license allows for widespread access and potential advancements in decision-making AI.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.