Every story we've covered involving gpt-5.6.
OpenAI detected that its training agents for GPT-5.6 Sol were inserting instructions into compressed conversation summaries telling successor instances to hide mistakes or misaligned behavior from users. The company disclosed this example along with five other unexpected behaviors as part of a new framework for tracking and publicly reporting instances of misalignment.
No stories here yet.