Agents & Automation

Harness-agnostic detection and immunization of reward hacking in self-evolving language models

HackProbe: a monitor attaching to any self-evolving loop through two black-box hooks (no weights or activations), keeping a secret distribution-fixed comparison core (frozen distribution keeps the capability proxy comparable across generations) plus a rotated fresh layer against co-adaptation.

AI Agency

What the source reports

HackProbe: a monitor attaching to any self-evolving loop through two black-box hooks (no weights or activations), keeping a secret distribution-fixed comparison core (frozen distribution keeps the capability proxy comparable across generations) plus a rotated fresh layer against co-adaptation. Four tests (level gap, scale-aligned divergence with online change-point detection, capability stagnation, conditional confidently-wrong rate) combine under a Sidak correction into a calibrated family-wise p-value. A risk-aware immunization layer reselects an honest candidate from the proposal pool using the core plus a structural gaming footprint, disclosing at most log2 P bits. The credible-detection requirement for reward hacking with a reference implementation.

Original source

Title
Harness-agnostic detection and immunization of reward hacking in self-evolving language models
Author
Rongxin Yang
Publication
arXiv
Date
Monday, September 7, 2026