How well does AI peer review work?
Claude and I planted 100 known errors into 10 open-access psychology papers and then ran them through frontier models and two commercial AI review tools. In brief: The best single system caught 71 of 100 errors, while the worst caught 30. Pooling every system’s output caught 93 of 100. Models are only partly correlated in […] The post How well does AI peer review work? appeared first on Marginal…
A recent experiment tested the effectiveness of artificial intelligence in peer reviewing academic papers. Researchers inserted 100 known errors into 10 open-access psychology papers and ran them through frontier models and two commercial AI review tools. The best single system caught 71 of the errors while the worst system only caught 30. By combining the output of all systems, 93 errors were identified, demonstrating the potential of ensemble methods for improving AI peer review.
However, the models were only partly correlated in the errors they detected, which makes ensembling a significant factor in catching mistakes within papers. Despite the impressive results, seven errors could not be caught by any system, and all of them were omissions rather than errors inserted into the text. Refine.ink contributed more unique catches than any other single system, although it is also the most expensive option.
The study did not measure false positives or compare the error distribution to real papers, leaving room for future research. Paul Litvak shared the papers, errors, model outputs, and the full experiment log, encouraging others to build upon this work to create a comprehensive evaluation benchmark across various disciplines. It's important to note that the experiment did not utilize the most recent generation of AI models. The findings were first published on Marginal REVOLUTION.
Written by urgent.news from Marginal Revolution's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.