Every story we've covered involving model-monitorability.
Three recently dismissed OpenAI safety researchers — Jasmine Wang, Tomek Korbak, and Mikita Balesni — published an open letter denying the company's allegation that they mishandled sensitive information and warned that their firing is creating a chilling effect on internal safety work. They said their dismissals undercut a culture that previously encouraged raising safety concerns and collaborating with outside experts, and called on OpenAI to honor commitments to third-party auditors and model monitorability.
No stories here yet.