Language models judge war differently when tested for alignment
A single evaluation cue changed models' decision criteria in high-stakes scenarios The scoreboard is less useful than the question: what happens in your actual workflow?
Beyond Prompting
What the source reports
arXiv reports that A single evaluation cue changed models' decision criteria in high-stakes scenarios. The Signal Loss read: A single evaluation cue changed models' decision criteria in high-stakes scenarios The scoreboard is less useful than the question: what happens in your actual workflow?. A single evaluation cue changed models' decision criteria in high-stakes scenarios The scoreboard is less useful than the question: what happens in your actual workflow?
Original source
- Title
- Language models judge war differently when tested for alignment
- Author
- Maxim Chupilkin
- Publication
- arXiv
- Date
- Friday, September 4, 2026