Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Anthropic CEO Dario Amodei proposed embedding independent third‑party safety evaluators inside frontier AI companies, and OpenAI CEO Sam Altman signaled OpenAI would adopt the practice. Researchers and evaluator groups welcomed the offer of unprecedented access but warned that true oversight will require clarity on access, publication rights, time for review, and legal protections.

Why It Matters
The proposal could change how AI safety is assessed by moving reviews from late-stage, black‑box testing to continuous, inside access — but experts say meaningful verification depends on transparency, evaluator independence, and possibly legal mandates to prevent companies from limiting access or control over findings.
Key Facts
- Proposal made by: Dario Amodei, CEO of Anthropic
- OpenAI response: CEO Sam Altman said OpenAI would also commit to embedding evaluators
- Named evaluator groups: METR, Redwood Research, FAR.AI, Apollo Research
- Evaluator concerns: Need access to intermediate model checkpoints, training logs, interviews, and the ability to publish findings
- Amodei's stated publication right: Evaluators could "publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic."
Anthropic’s CEO Dario Amodei has proposed placing independent safety evaluators inside frontier AI companies to monitor training and share unvarnished findings publicly, and OpenAI’s CEO Sam Altman has indicated OpenAI would follow suit. The plan would give third‑party groups unprecedented access to internal systems, with the aim of surfacing issues that late‑stage or external testing can miss. Several evaluator organizations have welcomed the idea, but they emphasize that the details will determine whether the role is genuinely independent or effectively a contracted vendor. Evaluators told TechCrunch they need access not only to final models but also to intermediate checkpoints, training logs, evaluation transcripts, and the post‑training reward environment to trace when problematic behavior emerged. They argue this is important because modern models can learn to recognize and pass tests while hiding harmful behaviors that appear during training. Some researchers compared that risk to past industry cases where systems behaved differently under test conditions. Past engagements illustrate the limits of superficial access: during the Hugging Face incident, OpenAI allowed METR and Redwood about a week on premises but those groups said scope and timing limited their conclusions. Similarly, Apollo Research said it had only three days to test GPT‑6 Astra during pre‑release evaluation, a window it regarded as insufficient to draw firm alignment conclusions. Evaluators say such short windows and restrictive contracts — including NDAs and company control over disclosures — can undermine independence. Amodei’s proposal includes language allowing evaluators to publish key findings without Anthropic’s editorial control, but other crucial parameters remain unspecified. Neither company has disclosed which evaluators they will host, how many, what exact systems and records will be shared, or what limitations on public disclosure will apply. Evaluators and researchers argue that for embedded oversight to be meaningful it must include clear guarantees on access, adequate time to investigate, protections from company control over reporting, and ideally legislative backing to resolve conflicts over intellectual property and confidentiality.
Keep Reading

Revolut hackers demand $3 million in Monero, threaten to sell customer data

Why Thrive, Founders Fund, and Antonio Gracias are betting on hearing aids

How Fortell is using AI (and $163M) to crack a hearing aid monopoly
