Scientists can’t stop using AI, even when forbidden from doing so
When scientists were asked to review papers for a conference without using AI tools, many of them ignored the request and used AI anyway
Researchers at Microsoft Research conducted a large-scale experiment during the 2026 International Conference on Machine Learning (ICML) conference held in Seoul, South Korea in July. The study involved 24,661 papers and 17,886 reviewers. Participants were offered two options for the review process - one strictly prohibiting the use of large language models (LLMs) while the other allowed LLMs for assistance, but not for evaluating the paper's merit or writing the review.
The results showed that papers reviewed under both policies had nearly identical acceptance rates of 27% and 26.5%, respectively. Average review scores were also very similar at 3.31 and 3.32 out of 6. Reviewer confidence levels remained unchanged regardless of the policy. However, reviews written under the permissive policy were longer by 5.5 to 7% and rated slightly higher in quality by experts, though no significant differences were detected when comparing reviews on a reviewer-by-reviewer basis.
An anonymous survey of 1,486 reviewers revealed that 22.5% of those instructed not to use AI admitted to doing so. The researchers primarily used the AI for brainstorming feedback, drafting review texts, reading papers, and summarizing their strengths and weaknesses. Miro Dudík, a Microsoft Research scientist and conference organizer, noted that the high reported non-compliance rate of 23% was higher than anticipated.
Kayvan Kousha from the University of Wolverhampton suggested that the temptation to use AI is partly due to the high reviewing workload faced by academics.
The study also utilized Pangram, an AI text detector, to analyze the reviews. Under the conservative policy, 52.2% of the reviews were classified as fully human-written, while this percentage dropped to 37.0% under the permissive policy. However, the researchers cautioned that AI text detectors are not infallible. Microsoft Research scientist Sunnie S. Y. Kim believes that their findings hold relevance for other computer science conferences and potentially other fields grappling with similar challenges related to AI usage.
Written by urgent.news from New Scientist's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.