AI judges need enough shared human checks
A study of ten AI judges finds that too little shared review by people can lead teams to choose the wrong judge or deploy it too soon. The authors test a way to plan which answers people should check. Their threshold is limited to clear-cut cases in this study, not student grading.
What the source reports
This paper tests ten AI systems used to judge other outputs across four sets of ratings. When too few items had both human and AI ratings, the study found wrong choices about which judge to use and whether to deploy it. The authors offer a formula for the minimum amount of shared review. They also spread human checks across groups of items to make those checks more useful. A 25 percent threshold is claimed only for cases that are not close calls under this study design. It is not a rule that certifies AI grading of student work.
Original source
- Title
- LLM Judge Validation Under Sparse Overlap: From Inference to Design
- Author
- Junxuan Li, Arko Mukherjee, Soumyabrata Pal
- Publication
- arXiv / NeurIPS 2026
- Date
- Friday, September 25, 2026