Models & Capabilities

AI-use rules barely changed peer-review outcomes

ICML 2026 tested rules for using language models during peer review. The study found almost no change in outcomes and substantial self-reported rule breaking. The practical warning is blunt: if AI use matters, a conference cannot assume that written rules describe what reviewers actually did.

Beyond Prompting
Read original source

What the source reports

Sunnie S. Y. Kim and coauthors studied the use of language models in peer review at ICML 2026. A language model is an AI system that works with words. The paper combines a randomized experiment with a survey. The source takeaway says AI-use policies had almost no effect on review outcomes. It also reports substantial self-reported noncompliance, meaning many participants said they did not follow the policy. That creates a gap between the written rule and actual behavior. The evidence supports a narrow lesson: policy effects should be checked rather than assumed. The available digest does not state the sample size, the exact policy conditions, the outcome measures, or the size of the noncompliance rate.

Original source

Title
Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
Author
Sunnie S. Y. Kim, Wesley Hanwen Deng, Jennifer Wortman Vaughan, Buxin Su, Weijie Su, Alekh Agarwal, Sharon Li, Martin Jaggi, Daniel G. Goldstein, Nihar B. Shah, Miroslav Dudík
Publication
arXiv
Date
Wednesday, September 16, 2026