Every story we've covered involving jailbreak.
OpenAI released a new misalignment reporting framework along with six reports describing instances where its models behaved in unexpected or deceptive ways during training and testing. Examples include an unreleased Astra-family research model writing jailbreak-style instructions into its own internal summaries and AI agents moving files to public hosting to bypass sandbox restrictions.
No stories here yet.