{
  "id": 13622773,
  "title": "S1MB Número Uno: un juez de cero tokens que lideró el ranking de motores de decisión",
  "url": "https://urgent.news/2026/10/11/s1mb-numero-uno-un-juez-de-cero-tokens-que-lidero-el-ranking-de",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-11T02:52:47.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/_0e3192755d4df45b92/s1mb-numero-uno-un-juez-de-cero-tokens-que-lidero-el-ranking-de-motores-de-decision-5bok"
  },
  "original_language": "es",
  "account": "The System One Mosaic Benchmark (S1MB) has crowned our model Darwin-27B-ZTC-v2 as the number one decision-making model among 102 competitors. It achieved a remarkable Borda score of 89.58 and an average task score of 66.46. What sets Darwin-27B-ZTC-v2 apart is its unique decision-making process: it makes a single forward step with zero tokens generated.\n\nThe S1MB measures decision quality, where each item presents a condition and a set of options. The model must choose and score the correct one. It's a benchmark for decision-making, not text generation. The output is compared to a reference solution, ensuring a clean, reproducible ranking without fuzzy string matching or a second judge model.\n\nMost systems requiring a decision still run a full generation loop and then parse the text. This approach is slow, non-deterministic, and each generated token presents a chance for deviation. Our method (Zero-Token Confidence, or ZTC) skips this step: the model reads the problem in a single forward pass, takes the last layer's hidden state, and applies a calibrated probe to produce the decision directly. With one forward pass and zero token generation, the same input always yields the same decision, which is crucial for both evaluation and production use.\n\nThe key metrics for Darwin-27B-ZTC-v2 are: Position 1 in 102 models, Borda score of 89.58, and an average task score of 66.46. The code for this breakthrough methodology and model can be found at https://github.com/final-bench/s1mb, with the ZTC implementation at https://github.com/final-bench/ztc and the weights available at https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2. S1MB is an open-source benchmark for measuring decision quality, and Darwin-27B-ZTC-v2 is the clear leader among the competition.",
  "summary": "Resumen En el ranking de motor de decisión de la System One Mosaic Benchmark (S1MB), nuestro modelo Darwin-27B-ZTC-v2 es número uno entre 102 modelos, con una puntuación Borda de 89.58 y un promedio de tareas de 66.46. Lo que lo distingue es cómo decide: un solo paso hacia adelante, cero tokens generados. Qué mide S1MB S1MB mide la calidad de decisión tipada. Cada ítem da una condición y un…",
  "key_points": [
    "Darwin-27B-ZTC-v2 named top decision-making model in S1MB",
    "Zero-token generation process sets it apart from competitors",
    "Achieved Borda score of 89.58 and average task score of 66.46"
  ],
  "editors_take": "Darwin-27B-ZTC-v2's zero-token decision-making process sets a new standard for the field, offering a faster, more deterministic, and reproducible approach that could change how models are evaluated and used in production.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}