Urgent.News

What's breaking now, across thousands of outlets.

Tech

A 25-verifier panel measured an effective size of 1.00

Generation got cheap. Trustworthy review did not. So we add reviewers. More eyes on the PR, more verifiers in the gate, a panel of LLM judges instead of one. The assumption underneath is that each additional reviewer adds independent evidence. That assumption is measurable. I measured it, and it did not hold. What IDKMesh is IDKMesh is an open-source research project (Apache-2.0, Python 3.11+)…

A research project called IDKMesh sought to understand how humans, AI agents, tools, and heterogeneous compute could work together on uncertain goals. The project measured the effectiveness of a panel of five verifier programs, each drawing inputs from a named region of a problem's input domain. These verifiers were designed to accept a candidate only if it matched a reference implementation on all of them.

The panel was found to be genuinely effective, with an average accuracy of 0.7956, and an effective size of 1.00 votes. However, when considering the correlation between the verifiers, the effective size was overstated by a factor of 1.66, making it no more reliable than a single verifier. This discrepancy highlights the limitations of relying on the standard correction method for estimating the effectiveness of a panel of reviewers.

The project's findings are based on a synthetic demonstration, not real-world evidence, and the verifiers used in the experiment are programs, not AI models. The repository is an open-source project under the Apache-2.0 license and can be cloned and run using Python.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Backups and other lies

Every sysadmin I know has the same secret: we do not actually back up half the things we tell other people to back up. I have been doing this for a decade.

More from Sunday 20 September →