Models & Capabilities

Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search

Environment-grounded audit of 12,249 self-reports across three model families: agents overstate top-100 success 4.8-9.3×; confidence uncalibrated, rationales barely consequential, selection doesn't improve reporting - self-reports are claims, not evidence

Beyond Prompting

What the source reports

Environment-grounded audit of 12,249 self-reports across three model families: agents overstate top-100 success 4.8-9.3×; confidence uncalibrated, rationales barely consequential, selection doesn't improve reporting - self-reports are claims, not evidence

Original source

Title
Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search
Author
Pan, Zhou & Hu
Publication
arXiv
Date
Tuesday, September 1, 2026