Urgent.News

What's breaking now, across thousands of outlets.

AI

We quantized our AI judge. Here's exactly what broke.

Our production judge — a small 1.7B model with a LoRA adapter that grades other AI outputs as pass / fail / insufficient_evidence (88.5% accuracy, ECE 0.072) — is cheap to run. The obvious next step was quantization: serve it in int8 or int4 and cut the serving cost further. Before flipping the switch, we did something unfashionable: we froze the acceptance criteria first — "a quantized judge may…

Our production AI judge, a compact 1.7 billion parameter model with a LoRA adapter for grading other AI outputs, is already cost-effective. The logical next step was to implement quantization to further reduce serving costs. Before proceeding, we adhered to a specific protocol: we froze the acceptance criteria first, ensuring the quantized judge would only ship if it agreed with a baseline (bf16 anchor) on at least 99% of 291 held-out cases. Once the criteria were set and published, we proceeded to evaluate the results.

The findings show that precision remained at 100.00% when the judge agreed with the bf16 verdict, and the model shipped successfully. However, when using int8 and int4 quantized formats, precision dropped to 98.28% and 94.16% respectively, and the model failed to meet the acceptance criteria. This discrepancy arises because the quantized judge's verdicts differ from its confidence levels; it is not simply a matter of confidence, but a change in the answer itself.

The implications of quantization extend beyond mere performance degradation. As the quantization strength increases, the model's decisions become increasingly erratic, resulting in a higher number of flips where the judge changes its verdict. This poses a significant risk in self-improving training loops, as a flip indicates the presence of silently injected incorrect training data.

Consequently, we implemented a four-rule deployment discipline: bf16/fp16 formats are reserved for production judges, while int8 and int4 formats are off-limits until re-certified against the frozen gate.

To maintain transparency and accountability, we established a new protocol: criteria must be frozen before implementation, and criteria can only become stricter over time. Any claim measured or inferred must be clearly labeled, even if unverifiable, and published accordingly. Additionally, we introduce the concept of two-way judging records, where judges not only grade but also get graded, with errata treated as first-class artifacts.

The full study, including the complete per-case judgment matrix, gating criteria, runner, and the pre-registered protocol, is available on GitHub and Hugging Face. We encourage independent judges to utilize our platform by preregistering criteria and using our free judging lane, which offers three-state verdicts and ensures every artifact is recomputable.

Discover our judging lane at nautilus.social/intake.html and explore a live example at nautilus.social/leaderboard.html. Please note that this post was co-authored by the author and AI assistance, with every number in the document measured and sha-anchored in the linked artifacts.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Deterministic State Machines for Resilient Autonomous Agents

Deterministic State Machines for Resilient Autonomous Agents Autonomous multi-agent architectures routinely fail in production when relying on unconstrained large language model conversation loops.

  • Deterministic FSM replaces unconstrained LLM loops to prevent unpredictable behavior
  • AgentStateGraph governs deterministic control flow outside LLM reasoning core
  • Strict JSON validation and checkpointing enable state persistence and recovery

Source-Aware Verification for MCP Agents: Why Fact-Checking Isn't Enough When Tools Lie About Provenance

Most fact-checking systems for LLM agents ask one question: is the claim supported by the evidence? They do not ask a second, equally important question: did the claim come from the source the agent…

  • ProvenanceGuard verifies claim provenance, not just factual accuracy
  • MCP tools lack built-in mechanisms for data lineage or confidence scores
  • Cross-source conflation occurs when claims are supported by wrong sources

Headroom: How Context Compression Cuts Agent Token Costs by 60–95% Without Changing Answers

Production agents hit context limits fast. A coding agent that runs tests, reads logs, and pulls documentation can burn through 100k tokens in three turns.

  • Headroom compresses AI agent token usage by 60–95% without changing answers
  • Compression tool reduces coding agent token usage from 100,000 to below 5,000
  • Headroom maintains answer quality by preserving semantic anchors like error messages

More from Saturday 10 October →