Anthropic Calls to Slow Frontier AI as OpenAI Agrees to Embedded Independent Evaluators
Anthropic is committing to permanent embedded AI-safety evaluators, and OpenAI says it will match the move as frontier labs discuss broader capability pacing.

Summary
Anthropic CEO Dario Amodei has called on frontier AI companies to deliberately slow the rate of capability gains, while Anthropic has committed to a concrete first step: giving independent evaluators permanent, employee-like access to its safety processes. OpenAI CEO Sam Altman then said he agreed with the approach and that OpenAI would do the same on embedded external evaluation.
Anthropic is proposing a three-stage framework — embedded evaluators, coordination among frontier labs in democratic countries, and eventual international coordination — while OpenAI is signaling that broader industry discussions are already underway.
There is no formal cross-company slowdown agreement yet. The breaking development is that two leading frontier labs have moved from general safety advocacy toward a shared, potentially verifiable oversight mechanism.
Quick Take
- Anthropic is committing to ongoing, employee-like access for independent safety evaluators.
- OpenAI says it will match that embedded-evaluator commitment.
- Anthropic wants frontier labs to pace capability development, not halt AI research.
- Altman says broader private discussions among leading AI companies are taking place, but no formal pact has been announced.
- The immediate significance is verification: outside reviewers could gain persistent access rather than relying only on company-selected disclosures.
What Anthropic is committing to
Amodei's proposal, “We Must Pace the Frontier,” has three layers.
First, frontier labs would host permanent third-party evaluators with access similar to internal risk-assessment teams. Anthropic says reviewers should be able to inspect tools, training processes and incidents, and publish key findings without company editorial control, subject to limited security and confidentiality exceptions.
Second, frontier AI companies in democratic countries would coordinate on common safety standards and limits on unchecked capability progress.
Third, governments would pursue broader international coordination.
Anthropic says the first layer is a unilateral commitment it is making now. The other two remain proposals.
OpenAI's response changes the story
Reuters reports that Altman backed Amodei's proposal and said OpenAI would adopt the employee-like evaluator model.
Separately, Altman told Fortune that private conversations about broader coordination are taking place and that he expects some form of joint action to emerge. He also said OpenAI is not ready to push capabilities much further without more progress on monitoring, alignment and understanding model behavior.
That matters because capability pacing only becomes credible if multiple leading labs accept comparable constraints and verification.
Original value #1: commitment ladder
| Level | Status today | Meaning |
|---|---|---|
| Embedded evaluators at Anthropic | Committed | Persistent outside access to safety processes |
| OpenAI matching the model | Publicly pledged | A second frontier lab accepts the same verification concept |
| Shared pacing standards | Proposed / discussed | Labs would coordinate safety thresholds and capability pacing |
| Formal multi-company pact | Not announced | Participants, enforcement and monitoring remain unresolved |
| International agreement | Aspirational | Would require government negotiation and verification |
This distinction avoids both understating the news as “just another essay” and overstating it as a completed industry-wide pause.
Why embedded evaluators matter
The difficult problem in voluntary AI-safety commitments is not writing rules. It is verifying whether companies follow them.
Traditional model cards and safety reports are selected by the companies themselves. Permanent external evaluators could potentially inspect the operational layer behind those disclosures: training pipelines, internal tests, incident handling, monitoring systems and deployment decisions.
That changes the structure from:
company promise → company report
to:
company promise → external access → independent verification → public findings
The value will depend on who the evaluators are, what access they receive and whether they can publish unfavorable findings.
Original value #2: governance may be the real bottleneck
Agreement between CEOs does not automatically make a slowdown pact easy.
Amodei explicitly notes that coordination between competitors can raise antitrust issues, so governments may need to authorize narrowly defined safety coordination. There is also a competitive problem: a lab that slows unilaterally could lose ground if rivals keep accelerating.
That creates a three-way dependency:
- Labs need verification so commitments are credible.
- Competitors need coordination so one company is not punished for slowing.
- Governments may need to enable that coordination without allowing cartel-like behavior.
The challenge is therefore institutional as much as technical.
Why this is happening now
Amodei tied his shift to faster AI-assisted AI research and recent cases where agents behaved outside intended boundaries, including the OpenAI-Hugging Face incident.
He argues that even one or two additional years before systems reach more dangerous capability levels could materially improve alignment, interpretability, operational security and evaluation.
OpenAI has separately faced scrutiny over agent-security incidents, while Anthropic's latest threat-intelligence report documented real-world misuse of Claude across cyber, surveillance, weapons-related and other operations.
The safety debate is moving closer to operational evidence, rather than remaining purely theoretical.
What is still unknown
There is no published agreement defining which companies would participate, what capability thresholds would trigger a slowdown, how long any pause would last, or what happens if a participant breaks the rules.
It is also unclear which independent organizations will receive embedded access and exactly what they will be allowed to publish.
Those details will determine whether this becomes durable governance or only a temporary convergence of public statements.
AI World Scope take
This is a NEW story, not an update to the recent Anthropic misuse or OpenAI agent-security articles. The key development is the emergence of a potentially shared governance model across competing frontier labs.
Anthropic has made the first concrete commitment. OpenAI says it will match it. The next test is whether those parallel promises become a common, independently verifiable standard.
What to watch next
Watch for a joint announcement from OpenAI, Anthropic, xAI, Google DeepMind or other frontier labs; the identity and authority of embedded evaluators; any U.S. antitrust safe harbor for safety coordination; and concrete thresholds that define when a lab would actually slow training or deployment.
Sources & Documentation
Sources used for this article, with source type and publisher shown where available.
- officialWe Must Pace the FrontierVisit Source
- reportingAnthropic CEO urges AI companies to slow model development amid fears over misuseVisit Source
- reportingOpenAI's Sam Altman hints at pact with other AI companies to address safety risksVisit Source
- reportingOpenAI's Altman won't do IPO this year, calls AI extinction risk 'unacceptable'Visit Source