Education & Public Impact

LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora

AI essay graders treated as raters: severity spans 219 points on ENEM's 0-1000 scale, judge-human correlations in the .47-.56 band, all five version contrasts shift severity (up to 133 points); self-consistent but not human-accurate

Covered by both lenses

What the source reports

AI essay graders treated as raters: severity spans 219 points on ENEM's 0-1000 scale, judge-human correlations in the .47-.56 band, all five version contrasts shift severity (up to 133 points); self-consistent but not human-accurate

Original source

Title
LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora
Author
arXiv, Veerendra Kumar Sunkavalli
Publication
arXiv
Date
Sunday, August 30, 2026

Also covered by the other newsletter.