AEF-1 Third-Party Evaluator Standard Draws xAI, OpenAI, and Anthropic Cosigners
Frontier labs are moving from fragmented safety tests toward shared external scrutiny as the AI Evaluator Forum publishes AEF-1, a baseline for independent third-party evaluations spanning access rules, conflict-of-interest controls, funding transparency, recusal, and disclosure. xAI, OpenAI, and Anthropic have cosigned the framework, while Anthropic separately pledged embedded third-party evaluators with desk-level access as pacing debates intensify. OpenAI meanwhile confirmed weeks of safety talks with Anthropic and Google DeepMind, saying antitrust waivers are not required, even as reporting highlights political pressure to prioritize speed over safeguards. The shift could give regulators and auditors a common comparison point across providers, yet published excerpts omit technical scoring criteria and leave unclear how binding any coordination will prove.
AEF-1 Third-Party Evaluator Standard Draws xAI, OpenAI, and Anthropic Cosigners
The AI Evaluator Forum published AEF-1, a proposed baseline for independent third-party AI evaluations covering access, conflicts of interest, funding relationships, recusal, and transparency. Anthropic is also unilaterally committing to embedded third-party evaluators with desk-level access as the pacing debate intensifies.
Key takeaway
AEF-1 cosignatures show major labs treating standardized third-party evaluation as competitive infrastructure, not optional PR.
What happened
The AI Evaluator Forum published AEF-1, a proposed baseline for independent third-party AI evaluations covering access, conflicts of interest, funding relationships, recusal, and transparency, according to Latent Space reporting. xAI, OpenAI, and Anthropic cosigned the standard, and Anthropic separately committed to embedded third-party evaluators with desk-level access as pacing debates intensify.
OpenAI confirmed it has been working with Anthropic and Google DeepMind for several weeks on AI safety steps and said it does not need an antitrust waiver to coordinate on safety, per Bloomberg Law reporting cited by Techmeme and Bloomberg Technology. TechCrunch reported the talks amid political pressure to keep pace with China, while separate coverage noted Anthropic calling to slow the AI race and OpenAI not going public this year.
Evidence
The AI Evaluator Forum published AEF-1 as a baseline for independent third-party AI evaluations.
Latent Space · attributed
The AI Evaluator Forum published AEF-1, a proposed baseline for independent third-party AI evaluations covering access, conflicts of interest, funding relationships, recusal, and transparency.
xAI, OpenAI, and Anthropic jointly endorsed the AEF-1 standard.
Latent Space · attributed
xAI, OpenAI, and Anthropic have jointly endorsed the AEF-1 standard, a unified protocol for third-party evaluators to assess AI model safety and performance.
Anthropic committed to embedded third-party evaluators with desk-level access.
Latent Space · attributed
Anthropic is also unilaterally committing to embedded third-party evaluators with desk-level access as the pacing debate intensifies.
OpenAI confirmed weeks of AI safety talks with Anthropic and Google DeepMind.
TechCrunch AI · attributed
OpenAI confirms weeks of AI safety talks with Anthropic and Google DeepMind, as Trump's team dismisses safety concerns and pushes to keep pace with China.
OpenAI said it does not need an antitrust waiver to coordinate on safety.
Techmeme · attributed
OpenAI says it has been working with Anthropic and Google for several weeks on AI safety, and it does not need an antitrust waiver to coordinate on safety
OpenAI is working with Anthropic and Google DeepMind on AI safety issues.
Bloomberg Technology · attributed
OpenAI is working on steps to address artificial intelligence safety issues with its top competitors Anthropic PBC and Google DeepMind, escalating industry efforts to respond to a groundswell of concern that the technology poses an economic and security threat.
AI leaders loosely agreed to pace development while political actors pushed different priorities.
Bloomberg Technology · attributed
AI leaders loosely agree to pace development, but China and Trump have something else in mind
Anthropic called for slowing the AI race and OpenAI will not go public this year.
Les Echos IA · attributed
Anthropic has publicly called for slowing down the AI race, while OpenAI has confirmed it will not go public this year.
A pause by major labs could affect chipmaker stock rallies tied to their hardware procurement.
Le Monde IA · attributed
Le Monde highlights the risk that if major AI labs like Anthropic, OpenAI, and xAI slow down their massive hardware procurement, the stock rally for chipmakers like Nvidia and AMD could stall.
OpenAI and Anthropic rivalry has fueled a race-like dynamic toward more powerful models.
Financial Times Technology · attributed
The deep schism between OpenAI and Anthropic has led to a race-like dynamic towards ever more powerful models
Why it matters
Regulators and researchers could compare models on a shared external baseline instead of relying on proprietary lab metrics, but voluntary talks may clash with political pressure to prioritize speed.
Limits and uncertainties
Published excerpts lack technical details on what AEF-1 measures, including red-teaming protocols, harm categories, or scoring methodologies.
Reporting does not specify what the cross-lab safety talks agreed to or whether commitments are binding rather than performative.
Practical implications
Builders should plan for more standardized third-party evaluation access requirements and transparency around funding and recusal.
Operators tracking frontier compute demand should watch whether pacing calls affect procurement cycles tied to chipmaker valuations.
What to watch
Whether AEF-1 publishes concrete evaluation criteria and which labs adopt embedded third-party evaluators with desk-level access.
Whether OpenAI, Anthropic, and Google DeepMind release enforceable safety coordination outcomes beyond voluntary statements.