A good AI score may still miss what a student got wrong
The study looked at student answers to college computer science questions. Three AI models gave scores that looked sound. Yet the models were less useful than teachers at finding mistaken ideas. A matching score is not the same as helpful feedback. This study does not test Texas schools.
What the source reports
Zhao, Jiao and Xu studied 3,041 student answers to 50 college computer science questions. They compared three commercial AI models with human teachers. The models gave rubric scores that lined up closely with one another. They were less useful than teachers at finding mistaken ideas in the answers. That gap matters if someone uses a score to plan what to teach next. This is a preprint about higher education. It does not test Texas K-12 lessons or prove that all AI scoring tools fail.
Original source
- Title
- Beyond Score Accuracy: Examining the Diagnostic Quality of LLM-Generated Structured Assessment in Higher Education
- Author
- Xi Zhao, Xinyue Jiao, Zhen Xu
- Publication
- arXiv
- Date
- Wednesday, October 7, 2026