Urgent.News

What's breaking now, across thousands of outlets.

AI

S1MB Número Uno: un juez de cero tokens que lideró el ranking de motores de decisión

Resumen En el ranking de motor de decisión de la System One Mosaic Benchmark (S1MB), nuestro modelo Darwin-27B-ZTC-v2 es número uno entre 102 modelos, con una puntuación Borda de 89.58 y un promedio de tareas de 66.46. Lo que lo distingue es cómo decide: un solo paso hacia adelante, cero tokens generados. Qué mide S1MB S1MB mide la calidad de decisión tipada. Cada ítem da una condición y un…

The System One Mosaic Benchmark (S1MB) has crowned our model Darwin-27B-ZTC-v2 as the number one decision-making model among 102 competitors. It achieved a remarkable Borda score of 89.58 and an average task score of 66.46. What sets Darwin-27B-ZTC-v2 apart is its unique decision-making process: it makes a single forward step with zero tokens generated.

The S1MB measures decision quality, where each item presents a condition and a set of options. The model must choose and score the correct one. It's a benchmark for decision-making, not text generation. The output is compared to a reference solution, ensuring a clean, reproducible ranking without fuzzy string matching or a second judge model.

Most systems requiring a decision still run a full generation loop and then parse the text. This approach is slow, non-deterministic, and each generated token presents a chance for deviation. Our method (Zero-Token Confidence, or ZTC) skips this step: the model reads the problem in a single forward pass, takes the last layer's hidden state, and applies a calibrated probe to produce the decision directly.

With one forward pass and zero token generation, the same input always yields the same decision, which is crucial for both evaluation and production use.

The key metrics for Darwin-27B-ZTC-v2 are: Position 1 in 102 models, Borda score of 89.58, and an average task score of 66.46. The code for this breakthrough methodology and model can be found at https://github.com/final-bench/s1mb, with the ZTC implementation at https://github.com/final-bench/ztc and the weights available at https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC-v2. S1MB is an open-source benchmark for measuring decision quality, and Darwin-27B-ZTC-v2 is the clear leader among the competition.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

One Self-Improving Model, Eleven Number-One Titles: What That Takes

TL;DR One self-improving model family now holds eleven public number-one benchmark records at the same time, across math, science, law, structured output, and decisions.

  • S1MB model breaks records in eleven categories simultaneously
  • Recursive self-improving loop pushes multiple benchmarks to top
  • Zero-token decision method enhances accurate and efficient decision-making

S1MB 1위: 0토큰으로 판정해 디시전 엔진 리더보드를 제패하다

요약 System One Mosaic Benchmark(S1MB) 디시전 엔진 리더보드에서 우리 모델 Darwin-27B-ZTC-v2가 102개 모델 중 1위입니다. Borda 점수 89.58, 태스크 평균 66.46. 핵심은 판정 방식입니다. 한 번의 순전파로, 생성 토큰 0개로 결정합니다. S1MB가 측정하는 것 S1MB는 타입드 결정 품질을 봅니다.

  • Darwin-27B-ZTC-v2 named top decision-making engine on S1MB leaderboard
  • Zero-Token Confidence method uses single forward pass, zero token generation
  • Model achieves 89.58 Borda score, 66.46 task average in S1MB benchmark

Engram Corruption: What Happens When a Skill Container Doesn't Own Its Payload

The setup In a modular AI framework like LivinGrimoire, behavior comes from small, swappable units called skills. A Brain holds them in lobes, and a skill's input() runs on every think cycle.

  • Skills are managed by higher-level AH skills, which handle skill management.
  • Engram snapshot process inadvertently includes payload skills, causing duplication issues.

I built Chloe: an open-source TypeScript framework for AI agents you control

I built Chloe for developers who want to build AI agents for businesses, especially small businesses. The idea is simple: use code for predictable tasks, and use AI where you need it.

  • I developed Chloe, an open-source TypeScript framework for AI agents.
  • Framework uses code for predictable tasks and AI for necessary operations.
  • Chloe provides developers control over agents, maintaining ownership of code.

More from Sunday 11 October →