Models & Capabilities
Model releases, capabilities, evaluations, system cards, and multimodal or agent performance.
-
Models & CapabilitiesRead story →
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Capability paradox: frontier LLM traders behave more alike as they get more capable; correlated behavior creates a non-diversifiable risk floor; when shared reasoning is accurate more agents help, when it's wrong the best models make the same mistake most efficiently - upgrading every agent to the newest model concentrates risk rather than reducing it
arXiv · Sep 3, 2026 -
Models & CapabilitiesRead story →
Claude Mythos 5.1 and Fable 5.1: Capabilities
Two "world's most powerful" models, six instruments, no consensus; Fable 5.1's FrontierCode scores drop at higher effort because it can't stop making helpful edits; Fable 5 stalled at ~11% of spend despite being "clearly best" - adoption gated by friction, not capability; Zvi's practice: dual-wield both models on non-trivial queries
Sep 5, 2026 -
Models & CapabilitiesRead story →
Claude Fable 5.1 and Mythos 5.1: The System Card
The card (Zvi Mowshowitz, "Claude Fable 5.1 and Mythos 5.1: The System Card," 09-04).
Don't Worry About the Vase · Sep 4, 2026 -
Models & CapabilitiesRead story →
Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search
Environment-grounded audit of 12,249 self-reports across three model families: agents overstate top-100 success 4.8-9.3×; confidence uncalibrated, rationales barely consequential, selection doesn't improve reporting - self-reports are claims, not evidence
arXiv · Sep 1, 2026 -
Models & CapabilitiesRead story →
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
The agent side (2609.03438, "Do GUI Agents Know When Not to Act?").
arXiv -
Models & CapabilitiesRead story →
Ugly AI Food Photos Are Only the Beginning
Reece Rogers on Ugly AI Food Photos Are Only the Beginning. Relevant for anyone choosing or evaluating frontier models.
WIRED · Sep 3, 2026 -
Models & CapabilitiesRead story →
Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence
The epistemic claim (2609.01873, "Epistemic Sybil Resistance"): another agent is not another observation.
arXiv