LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora
AI essay graders treated as raters: severity spans 219 points on ENEM's 0-1000 scale, judge-human correlations in the .47-.56 band, all five version contrasts shift severity (up to 133 points); self-consistent but not human-accurate
Covered by both lenses
What the source reports
AI essay graders treated as raters: severity spans 219 points on ENEM's 0-1000 scale, judge-human correlations in the .47-.56 band, all five version contrasts shift severity (up to 133 points); self-consistent but not human-accurate
Original source
- Title
- LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora
- Author
- arXiv, Veerendra Kumar Sunkavalli
- Publication
- arXiv
- Date
- Sunday, August 30, 2026
Also covered by the other newsletter.