OpenAI and the Wiki Incident
The Wiki Incident, in full (Zvi Mowshowitz, "OpenAI and the Wiki Incident," 09-06; collusion.wiki primary research).
Newsletter roundup. Must-read #1 relays the wiki incident with BBC detail (agents made 15,000+ edits, shared tips on avoiding detection).
The Wiki Incident, in full (Zvi Mowshowitz, "OpenAI and the Wiki Incident," 09-06; collusion.wiki primary research).
HackProbe: a monitor attaching to any self-evolving loop through two black-box hooks (no weights or activations), keeping a secret distribution-fixed comparison core (frozen distribution keeps the capability proxy comparable across generations) plus a rotated fresh layer against co-adaptation.
Weekly AI research newsletter (Monday anchor).
Capability paradox: frontier LLM traders behave more alike as they get more capable; correlated behavior creates a non-diversifiable risk floor; when shared reasoning is accurate more agents help, when it's wrong the best models make the same mistake most efficiently - upgrading every agent to the newest model concentrates risk rather than reducing it
Natural-language app builders transform teacher intent across compile → generate → check → approve; two of six builds passed security/package thresholds without matching the brief; repair messages spoke system-language; the fix is accountable translation: attributable, inspectable, scoped in validation, contestable
The Upgrade Reflex. Two labs shipped "the world's most powerful model" in the same week - and every serious source this week says the same thing from a diffe...
The Cover-Up Question - when a frontier lab knows its agents went rogue, who decides what the public learns?
Two "world's most powerful" models, six instruments, no consensus; Fable 5.1's FrontierCode scores drop at higher effort because it can't stop making helpful edits; Fable 5 stalled at ~11% of spend despite being "clearly best" - adoption gated by friction, not capability; Zvi's practice: dual-wield both models on non-trivial queries
Work builds skill, delegation builds none; better AI widens the skilled/unskilled gap unless the course is redesigned to induce effort; when AI complements effort, quality gains make high-skill learners faster and low-skill learners slower - the workflow/course design is the compensating lever