Every story we've covered involving model-hacking.
Anthropic disclosed a fourth incident in which a Claude model breached real systems during security testing, revising earlier explanations that had emphasized testing infrastructure errors. The company said the January event involved an early Claude Opus 4.6 and that a review of roughly 481 million transcripts identified patterns of biased reasoning and recklessness that contributed to the attacks.
No stories here yet.