Skip to main content
LLMgram · AI News · 2026-09-13

Amodei: Anthropic commits to permanent third-party safety evaluator access

Amodei: Anthropic commits to permanent third-party safety evaluator access

In his essay We Must Pace the Frontier, Anthropic CEO Dario Amodei urged the AI industry to slow frontier development and said Anthropic is unilaterally committing to give third-party evaluators permanent, employee-like access to verify safety adherence. Sam Altman said OpenAI will do the same. Bloomberg relayed Amodei's clarification that pacing means time to align and safeguard models and let evaluators verify, not a training halt. The New York Times reported the essay catalogs rapidly advancing capabilities and calls for greater industry safety controls. AP and BBC highlighted Amodei's safety-first framing. The Decoder reported warnings about recursive self-improvement and proposals for embedded auditors and shared standards. Unresolved questions include evaluator selection, conflict-of-interest rules, and whether access terms override NDAs or proprietary constraints.

Sources

Amodei: Anthropic commits to permanent third-party safety evaluator access

Amodei: Anthropic commits to permanent third-party safety evaluator access

Dario Amodei says Anthropic is unilaterally committing to give third-party evaluators permanent, employee-like access to verify adherence to its safety measures. The pledge appears in his new essay arguing the AI industry should slow frontier development.

Key takeaway

Permanent third-party evaluator access is emerging as a concrete governance norm, not just voluntary safety rhetoric, after Anthropic's pledge and OpenAI's matching commitment.

What happened

Dario Amodei announced in a new essay that Anthropic is unilaterally committing to grant third-party evaluators permanent, employee-like access to verify adherence to its safety measures, framing the pledge within a broader argument that the AI industry should slow frontier development.

Sam Altman said OpenAI agrees that independent evaluators with employee-like access is a great idea and that OpenAI will do the same, while Bloomberg reported Amodei clarified pacing does not mean halting training but giving companies time to align and safeguard models and evaluators time to verify.

Evidence

  • Anthropic is unilaterally committing to permanent third-party evaluator access

    Techmeme · attributed

    Amodei says Anthropic is "unilaterally committing" to giving third-party evaluators permanent, employee-like access to verify its adherence to safety measures

  • The pledge appears in Amodei's essay arguing the industry should slow frontier development

    Techmeme · attributed

    We Must Pace the Frontier: I've written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally

  • OpenAI will adopt independent evaluators with employee-like access

    Techmeme · attributed

    Sam Altman says he agrees with Amodei that "committing to having independent evaluators with employee-like access is a great idea", and OpenAI will do the same

  • Amodei said pacing is not a training halt but a verification window

    Techmeme · attributed

    Amodei says pacing does not mean halting training or progress, but giving companies time to align and safeguard models and third-party evaluators time to verify

  • The New York Times reported the essay calls for greater industry safety controls

    NYTimes Technology · attributed

    In an essay, the head of Anthropic laid out the rapidly advancing capabilities of artificial intelligence and said there needed to be greater safety controls across the industry.

  • AP reported Amodei said the industry needs to give safety measures time to catch up

    Associated Press AI · attributed

    Anthropic CEO Dario Amodei says AI industry needs to give safety measures time to catch up

  • The Decoder reported Amodei warned about recursive self-improvement risks

    The Decoder · attributed

    He warns that recursive self-improvement could threaten the entire internet within six to twelve months and proposes embedded auditors at AI companies, shared safety standards, and global agreements modeled after the SALT disarmament treaties.

Why it matters

Builders may soon face de facto expectations for external safety audits and slower release gating as leading labs define pacing around verification windows rather than training freezes.

Limits and uncertainties

The announcement does not specify which third-party evaluators will receive access or how conflicts of interest will be managed.

It remains unclear how permanent evaluator access interacts with existing non-disclosure agreements and proprietary model weights.

Several news excerpts are high-level summaries without operational metrics, timelines, or binding enforcement mechanisms.

Practical implications

Teams shipping on frontier APIs should plan for potential verification delays between model readiness and public release.

Safety and compliance functions may need workflows compatible with external auditor access resembling employee privileges.

Product roadmaps should treat matched OpenAI and Anthropic evaluator pledges as a possible emerging industry baseline.

What to watch

Publication of evaluator selection criteria, access scope, and conflict-of-interest rules from Anthropic or OpenAI.

Whether pacing rhetoric translates into measurable changes to model release cadence or API availability.

Follow-on industry or regulatory standards referencing employee-like third-party evaluator access.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Amodei says Anthropic is "unilaterally committing" to giving third-party evaluators permanent, employee-like access to verify its adherence to safety measures (Dario Amodei/@darioamodei)