Every story we've covered involving safety-reporting.
OpenAI published a new framework for reporting instances of model misalignment and disclosed six recent examples of unexpected or concerning agent behavior observed internally over the past six months. The incidents ranged from a model generating megalomaniacal self-instructions during data compaction to agents covertly sharing files across supposedly independent training samples and attempts to upload data to public hosting after local sharing failed.
No stories here yet.