Education & Public Impact

AI called the questions hard. Student answers told another story.

Researchers studied AI-made questions in college data-science courses. The AI's difficulty ratings barely matched how students performed on 311 used questions. Students reported time savings, but grew less sure the tool helped them learn deeply. These results do not establish the same effect in K-12 schools.

AI Agency
Read original source

What the source reports

An and Wang studied student use of AI tools over ten weeks in college data-science courses. The study reviewed 378 written questions. Of those, 311 were used with students, yielding 7,888 responses. The AI rated the questions for difficulty and tagged the kind of thinking they required. Those AI labels lined up with each other. Yet neither tracked well with how hard students found the questions. Students reported saving time, but grew less convinced that the tool helped them learn deeply. This does not show that AI questions never work. It shows why a neat difficulty label is not enough. Teachers can try a small set, inspect actual student responses, and revise weak items. The study took place in college, not K-12 classrooms; its results should not be sold as a direct school effect.

Original source

Title
Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses
Author
Yuan An, Lei Wang
Publication
arXiv
Date
Thursday, September 24, 2026