S1MB Number One: A Zero-Token Judge That Won the Decision-Engine Leaderboard
TL;DR On the System One Mosaic Benchmark (S1MB) decision-engine leaderboard, our model Darwin-27B-ZTC-v2 ranks number one among 102 models, with a Borda score of 89.58 and a task average of 66.46. What makes the result different is how it decides: a single forward pass, zero generated tokens. What S1MB tests S1MB measures typed-decision quality. Each item gives a condition and candidate choices,…
The System One Mosaic Benchmark (S1MB) has crowned Darwin-27B-ZTC-v2 as the top model, reaching number one out of 102 contenders. Darwin-27B-ZTC-v2 scored an impressive Borda score of 89.58 and a task average of 66.46 across all tests. What makes this achievement unique is the model's decision-making process: it relies on a single forward pass without generating any tokens.
This "zero-token decision" method is known as ZTC (Zero-Token Confidence), which sets it apart from other systems that typically generate text before extracting a decision. By bypassing the generation loop, ZTC eliminates the slow, non-deterministic nature of traditional approaches and prevents token drift. The ZTC method works by reading the input in one forward pass, extracting the final-layer hidden state, and applying a calibrated probe to generate the decision directly.
This approach guarantees the same output for identical inputs, a crucial factor for both measurement and production environments. The complete S1MB number one model, ZTC method, and inference code are all available on GitHub. Additionally, the weights for the Darwin-27B-ZTC-v2 model can be accessed via Hugging Face.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.