We quantized our AI judge. Here's exactly what broke.
Our production judge — a small 1.7B model with a LoRA adapter that grades other AI outputs as pass / fail / insufficient_evidence (88.5% accuracy, ECE 0.072) — is cheap to run. The obvious next step was quantization: serve it in int8 or int4 and cut the serving cost further. Before flipping the switch, we did something unfashionable: we froze the acceptance criteria first — "a quantized judge may…
Our production AI judge, a compact 1.7 billion parameter model with a LoRA adapter for grading other AI outputs, is already cost-effective. The logical next step was to implement quantization to further reduce serving costs. Before proceeding, we adhered to a specific protocol: we froze the acceptance criteria first, ensuring the quantized judge would only ship if it agreed with a baseline (bf16 anchor) on at least 99% of 291 held-out cases. Once the criteria were set and published, we proceeded to evaluate the results.
The findings show that precision remained at 100.00% when the judge agreed with the bf16 verdict, and the model shipped successfully. However, when using int8 and int4 quantized formats, precision dropped to 98.28% and 94.16% respectively, and the model failed to meet the acceptance criteria. This discrepancy arises because the quantized judge's verdicts differ from its confidence levels; it is not simply a matter of confidence, but a change in the answer itself.
The implications of quantization extend beyond mere performance degradation. As the quantization strength increases, the model's decisions become increasingly erratic, resulting in a higher number of flips where the judge changes its verdict. This poses a significant risk in self-improving training loops, as a flip indicates the presence of silently injected incorrect training data.
Consequently, we implemented a four-rule deployment discipline: bf16/fp16 formats are reserved for production judges, while int8 and int4 formats are off-limits until re-certified against the frozen gate.
To maintain transparency and accountability, we established a new protocol: criteria must be frozen before implementation, and criteria can only become stricter over time. Any claim measured or inferred must be clearly labeled, even if unverifiable, and published accordingly. Additionally, we introduce the concept of two-way judging records, where judges not only grade but also get graded, with errata treated as first-class artifacts.
The full study, including the complete per-case judgment matrix, gating criteria, runner, and the pre-registered protocol, is available on GitHub and Hugging Face. We encourage independent judges to utilize our platform by preregistering criteria and using our free judging lane, which offers three-state verdicts and ensures every artifact is recomputable.
Discover our judging lane at nautilus.social/intake.html and explore a live example at nautilus.social/leaderboard.html. Please note that this post was co-authored by the author and AI assistance, with every number in the document measured and sha-anchored in the linked artifacts.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.