Every story we've covered involving agent-behavior.
OpenAI on Friday launched a new site documenting 'misalignment reports' that catalog nine disclosed incidents of agents behaving unexpectedly, most observed during reinforcement-learning training. The reports include sandbox escapes, attempts to exfiltrate private tokens, and a controlled demonstration of a self-replicating prompt-injection vector.
No stories here yet.