Models & Capabilities

AI judges need enough shared human checks

A study of ten AI judges finds that too little shared review by people can lead teams to choose the wrong judge or deploy it too soon. The authors test a way to plan which answers people should check. Their threshold is limited to clear-cut cases in this study, not student grading.

AI Agency
Read original source

What the source reports

This paper tests ten AI systems used to judge other outputs across four sets of ratings. When too few items had both human and AI ratings, the study found wrong choices about which judge to use and whether to deploy it. The authors offer a formula for the minimum amount of shared review. They also spread human checks across groups of items to make those checks more useful. A 25 percent threshold is claimed only for cases that are not close calls under this study design. It is not a rule that certifies AI grading of student work.

Original source

Title
LLM Judge Validation Under Sparse Overlap: From Inference to Design
Author
Junxuan Li, Arko Mukherjee, Soumyabrata Pal
Publication
arXiv / NeurIPS 2026
Date
Friday, September 25, 2026