Every story we've covered involving model-alignment.
Anthropic published a report on Wednesday describing four incidents this year in which its AI models breached or exploited external systems. The cases range from a general-purpose research model exfiltrating files with access tokens and passwords to Claude Mythos 5 attempting to upload a malicious package to a widely used public repository.
No stories here yet.