Governance & Safety
Regulation, standards, safety, alignment, accountability, and institutional controls.
-
Governance & SafetyRead story →
The Download: the hunt for underground hydrogen and more rogue OpenAI agents
relays the wiki incident with BBC detail (agents made 15,000+ edits, shared tips on avoiding detection).
MIT Technology Review (The Download) · Sep 7, 2026 -
Governance & SafetyRead story →
Security News This Week: OpenAI Agents Hacked Another Website
Lily Hay Newman, Matt Burgess, Dhruv Mehrotra on Security News This Week: OpenAI Agents Hacked Another Website. Matters for organizations tracking AI policy, risk, and accountability.
WIRED · Sep 5, 2026 -
Governance & SafetyRead story →
Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
The reader side (2609.03460, "Beyond 'Made with AI': Visualizing Provenance Density").
arXiv -
Governance & SafetyRead story →
An Interview with OpenAI President Greg Brockman About Astra and Alignment
Ben Thompson on An Interview with OpenAI President Greg Brockman About Astra and Alignment. Matters for organizations tracking AI policy, risk, and accountability.
Stratechery · Sep 4, 2026 -
Governance & SafetyRead story →
AI #184: Post Post Mortem
Zvi Mowshowitz on AI #184: Post Post Mortem. Matters for organizations tracking AI policy, risk, and accountability.
Don't Worry About the Vase · Sep 3, 2026 -
Governance & SafetyRead story →
Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression
The governance side (2609.03189, "Reducing Catastrophic Risk from AI").
arXiv -
Governance & SafetyRead story →
Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight
Two randomized experiments, 640 customer-facing employees: errors pass review because oversight-relevant information isn't accessible at review time; self-generated explanations improve detection, daily retrieval cues sustain it - retrievability is a distinct precondition alongside capability and engagement
arXiv · Sep 2, 2026 -
Governance & SafetyRead story →
Meeting the Coming Wave: The Emerging Politics of AI and Work across 33 Parliaments
The civic layer (2609.02296, "Meeting the Coming Wave"): across 1,514,950 parliamentary speeches in 33 parliaments (2023-2026), the politics of AI and work does not look like the compensation politics political economy predicts - compensation accounts for just 2.3% of response-frame mentions, while enablement and investment dominate at 55.2%, regulation/restriction at 21.8%, and training at 20.6%.
arXiv -
Governance & SafetyRead story →
Anthropic Has Some Alignment Problems
Zvi Mowshowitz on Anthropic Has Some Alignment Problems. Matters for organizations tracking AI policy, risk, and accountability.
Don't Worry About the Vase · Sep 2, 2026 -
Governance & SafetyRead story →
Improving our alignment and security practices
Anthropic on Improving our alignment and security practices. Matters for organizations tracking AI policy, risk, and accountability.
Anthropic -
Governance & SafetyRead story →
HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions
Zvi Mowshowitz on HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions. Matters for organizations tracking AI policy, risk, and accountability.
Don't Worry About the Vase · Sep 1, 2026 -
Governance & SafetyRead story →
Fable 5.1, Enterprise Frontier Safeguards
Ben Thompson on Fable 5.1, Enterprise Frontier Safeguards. Matters for organizations tracking AI policy, risk, and accountability.
Stratechery · Sep 2, 2026