
Summary
Anthropic CEO Dario Amodei has called on the AI industry to slow the pace of frontier-model capability gains and give independent evaluators employee-like access to safety work. OpenAI says it will adopt a similar commitment, while the proposal raises unresolved questions about antitrust coordination, regulatory standards and enforceability.
A call to slow capability gains, not stop research
Anthropic CEO Dario Amodei has urged the artificial-intelligence industry to slow the rate at which it improves the capabilities of frontier models. His proposal does not call for ending research or freezing development. Instead, he argues that labs should leave enough time to align, test and safeguard increasingly powerful systems before pushing them further.
“Pacing the frontier,” as the idea has been described, is therefore less a technical limit than a governance proposal. Amodei says progress can remain fast while companies use the time gained to address safety concerns. The intervention comes as the largest AI labs compete to build systems that can reason, use tools and help construct the next generation of models.
The argument is partly based on the pace of recent progress. According to reporting on Amodei’s post, he believes AI capabilities have advanced particularly quickly in recent months, including the ability of models to contribute to the development of future AI systems. That prospect creates a feedback loop: more capable systems may not simply perform tasks better, but may also accelerate the work that produces even more capable systems.
Amodei also pointed to the attack involving Hugging Face and to the possibility of multiple AI agents acting together. He warned that, within six to twelve months, a coordinated swarm could potentially take over parts of the internet through a persistent botnet, with economic damage reaching hundreds of billions of dollars. The reporting does not establish that such a scenario has occurred. It presents it as a risk assessment underlying Anthropic’s call for a more cautious pace.
Anthropic’s embedded-evaluator proposal
The most concrete element of Amodei’s plan is the use of “embedded evaluators.” Anthropic says it is unilaterally committing to allow outside teams to assess its systems from inside the company, with desks, badges, company laptops and permissions close to those available to internal risk personnel. The evaluators would also be free to publish what they find.
That approach is more intrusive than a conventional external audit. A periodic audit may be limited to selected documentation, predefined tests or a narrow stage of the development cycle. An embedded team, by contrast, could observe how safety work is conducted over time and examine the relationship between a model’s capabilities and its safeguards. In principle, that could provide earlier warning about risks involving autonomous operation, cyber activity, code generation, internet access or coordination among agents.
The proposal also creates difficult questions about independence. Who chooses the evaluators? Who pays them? What information may they access? How are trade secrets and sensitive security details protected? If an evaluator identifies a serious problem, can it require a release to be delayed, or can it only publish a report after the fact? The source material does not answer these questions.
Allowing evaluators to publish findings may strengthen accountability, but it can also create disclosure and liability issues. A report could reveal exploitable weaknesses, expose confidential research or create disputes over whether a risk was adequately mitigated. For the model to work, access rules, publication procedures and escalation powers would need to be defined in advance. A public commitment alone does not establish those safeguards.
OpenAI backs the idea, but a shared standard is still missing
OpenAI CEO Sam Altman said he agreed with Amodei that the frontier needs to be paced. He described independent evaluators with employee-like access as a strong idea and said OpenAI would make the same commitment, adding that pacing had been a primary topic of internal discussions in recent weeks. Elon Musk, who leads xAI, posted a brief expression of support.
The public response means that two major frontier labs have endorsed a similar principle within a short period. That is significant, but it does not yet amount to a common industry standard. Companies could define “independent,” “employee-like access” and “safety evaluation” in materially different ways. Without comparable protocols, outside observers may still be unable to determine whether one lab’s safeguards are stronger than another’s.
The issue is particularly important for organizations that may eventually use advanced models in sensitive operational environments. If an AI system assists with access controls, compliance review, fraud monitoring or other high-consequence workflows, a general statement that the model has been evaluated may not be enough. Institutions need to know which capabilities were tested, under what conditions, by whom and with what authority to disclose problems.
Why Amodei wants an antitrust pathway
Amodei’s second major request concerns competition law. He has asked the US government to mediate or narrowly permit safety discussions among competitors. His concern is that companies may be reluctant to share information about threats, testing methods or incidents if coordination itself could create antitrust exposure.
The proposal is not a request for unrestricted cooperation. Any safe harbor would need to distinguish legitimate safety coordination from agreements that limit competition, divide markets or align commercial conduct. A narrowly drafted mechanism might specify who can participate, which subjects can be discussed, what information can be exchanged and how long the arrangement lasts. It might also require records or government review.
That distinction is central. Joint work on safety can produce public benefits, especially when companies face common technical threats. At the same time, the largest AI labs already hold substantial influence over model development, infrastructure and distribution. A broad exemption could make it harder for regulators to detect coordination that goes beyond safety. Amodei’s proposal therefore shifts part of the debate from technical evaluation to institutional design.
Europe offers a different regulatory reference point
The reporting also highlights a gap in the framing of the proposal: Amodei’s post does not mention Europe, even though the European Union has already adopted obligations for general-purpose models with systemic risk. Under the EU AI Act rules described in the source material, providers must carry out documented adversarial testing using standardized protocols, assess and mitigate systemic risks at Union level, and protect models against cybersecurity threats.
Those rules do not make Anthropic’s proposed mechanism unnecessary, nor do they prove that legal compliance is equivalent to robust safety. Regulation establishes duties, documentation and enforcement responsibilities. Embedded evaluators emphasize continuous scrutiny, operational access and the ability to publish findings. The two approaches may complement one another, but they can also differ over testing scope, disclosure and the authority to intervene.
The European comparison also complicates the antitrust request. The EU no longer relies on the old model of granting individual exemptions in the way it once did, and the competition-law treatment of cooperation depends on the specific conduct and its effects. A US-style narrow waiver would not automatically translate into a European solution. Any cross-border safety framework would need to account for different competition regimes as well as different approaches to AI oversight.
The real test is verifiability
The importance of the Anthropic and OpenAI announcements lies in their attempt to turn broad safety warnings into organizational commitments. The debate is now about who can inspect a model, what access they receive, whether findings can be made public and how competitors can share information without weakening competition.
For institutional wallet and custody operations, the issue is relevant only insofar as increasingly capable models may be integrated into authorization, monitoring or compliance workflows. In such environments, the central question is not whether an AI provider has made a safety pledge, but whether the provider can demonstrate what was tested, what failed, how failures were addressed and who had the authority to challenge deployment decisions.
The public proposals still lack a detailed timetable, common metrics, a clearly independent oversight body and defined consequences for noncompliance. They also face a basic commercial tension: frontier labs gain competitive advantage from rapid iteration, while deeper testing and external access can slow releases and increase costs.
The next stage will require clearer answers to at least three questions. What capabilities should trigger enhanced review? What evidence is sufficient to show that a systemic risk has been reduced? Where is the boundary between safety cooperation and unlawful coordination? Until those questions are answered, the latest statements are best understood as signals of a possible governance direction, not as completed safety regimes.
Source: link