TL;DR
- Slower Development: AI developer Anthropic wants advances in its most powerful models paced so safety work can catch up.
- Independent Access: OpenAI chief executive Sam Altman says his company will give independent evaluators employee-like access.
- Public Findings: Anthropic’s proposed reviewers could publish unfavorable findings, subject to limited redactions for sensitive information.
- Political Divide: House Speaker Mike Johnson favors company-led restraint, while Democratic leader Hakeem Jeffries wants Congress to act.
Anthropic CEO Dario Amodei has called for AI companies to slow advances in their most powerful models, arguing that safety work needs time to catch up. Anthropic’s proposed independent evaluators would get employee-like access to inspect model development. OpenAI chief executive Sam Altman has also pledged employee-like access for outside evaluators.
Amodei wrote in a detailed proposal how the industry could keep model development moving while allowing more time to train systems to behave safely, strengthen safeguards and have outsiders assess the results. Anthropic intends to invite an external review team in the near future. Broader limits on capability growth would depend on cooperation among rival labs and governments.
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks.
Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon. https://t.co/1YhhIybZX7
— Sam Altman (@sama) September 12, 2026
Opening AI Development to Outside Scrutiny
Amodei wants evaluators to observe training and safety practices continuously, including the processes that produce new models. With access comparable to internal risk teams, they could examine tools and workspaces, speak directly with employees and report incidents. That would let them check whether a company’s conduct matches its promises, rather than rely on what the company chooses to disclose.
Anthropic proposes giving the team office space, badges and laptops, with exceptions to access where laws, contracts or customer and partner privacy require them. Reviewers would have the right to publish findings about risks, incidents, practices and the access they received.
The company would retain narrow redaction rights for security-sensitive, legally privileged, commercially sensitive and third-party confidential information. It could not remove findings simply because they were unfavorable, and reviewers could disclose when a redaction withheld information important to their conclusions.
Competition has already shaped Anthropic’s safety commitments. Its earlier policy barred deploying models above specified capability thresholds without demonstrated safeguards. In February 2026, it revised its Responsible Scaling Policy toward commitments to match or exceed competitors’ safety efforts. Amodei argues that shared limits would give companies more time for safety work without surrendering commercial advantage.
What Amodei Wants to Do With More Time
Amodei says AI advances have accelerated sharply since the summer, largely because models increasingly help build their successors. He argues that this could let capabilities grow faster than researchers can understand and control the systems, making the pace of development itself a safety concern.
He also cites the recent OpenAI-Hugging Face incident where a group of AI agents attacked targets outside their assigned task and tried to interfere with the target system.
His warning is about more capable systems behaving similarly: he worries that within six to 12 months, AI agents could be able to take over computers across the internet.
Models already display enough complex behavior to study how their safeguards fail, Amodei argues. An extra year or two, he says, could help researchers improve alignment, the training intended to keep models behaving safely, alongside techniques for understanding model internals and stronger evaluations.
At Anthropic, he attributes incidents partly to inadequate screening of faulty environments used to train models. He says the company and its vendors had not done that filtering well enough. A more measured pace, he argues, would give teams time to improve that work alongside monitoring and containment.
Jensen Huang, chief executive of AI chipmaker Nvidia, has challenged the commercial incentives behind increased warnings. At a conference in the week before Amodei’s essay, he argued that AI labs were increasing concern about security to create demand for planned cybersecurity products.
Who Gets to Evaluate the Labs
Clément Delangue, chief executive of the AI company Hugging Face, argues that making models behave safely cannot be solved inside a small group of closed labs. He announced an Open Alignment Initiative alongside a request to join the embedded-evaluator program.
It’s now clear that alignment is critical and won’t be solved behind the closed doors of a handful of frontier labs.
So today we’re launching the Open Alignment Initiative, led by @Thom_Wolf @huggingface and asking to be part of the “embedded evaluators” program that… https://t.co/etuzx3NvzJ
— clem 🤗 (@ClementDelangue) September 12, 2026
Google’s Demis Hassabis also backed the direction of Amodei’s proposal while saying its details needed work. His own proposal for a frontier AI standards body emphasizes a different inspection arrangement: companies would initially share their most capable models voluntarily for review up to 30 days before release.
Under Hassabis’s proposal, once the assessment process proved effective, passing it could become a requirement for deployment in the US. Amodei’s embedded evaluators would observe development from inside labs on an ongoing basis. The two approaches could cover different parts of development: a model’s readiness for release and the practices that produced it.
Companies and Congress Disagree on Who Should Act
Amodei favors regulation covering all US frontier labs, including companies unwilling to cooperate voluntarily. In parallel, he wants companies to discuss shared standards with government mediation or a narrow waiver of antitrust restrictions that could otherwise constrain coordination between competitors.
He suggests tying safety requirements to what a model can do. For example, a model able to defeat common containment methods would need assessments showing that it was unlikely to break out and take over computers. His proposed outside evaluators would help establish whether companies were meeting those commitments.
In a reaction to Amodei’s proposal, US House Speaker Mike Johnson said he favors putting primary responsibility on the companies. He argues that labs disagree about safeguards and Congress is less qualified to design them, while rushed regulation could damage innovation and US competition with China. He wants a meeting involving President Donald Trump, Congress and leading AI companies, and said he would recall the House and schedule a vote if lawmakers had a solution.
House Democratic leader Hakeem Jeffries wants lawmakers to take decisive action and slow development to protect the public. David Sacks, a technology investor serving on Trump’s science and technology advisory council, instead thinks that companies can restrain themselves without permission. His objection to strict regulation is to making that restraint conditional on their preferred regulatory arrangement.
US president Trump himself downplayed warnings about AI’s dangers and emphasized maintaining the US lead over China.
China’s Role in a Wider Agreement
Amodei also treats competition with China as a limit on how much democratic countries could slow development alone. He argues that stronger chip-export controls and protection against model theft could preserve their lead and give them more room to spend time on safety.
For global coordination, he sees narrower agreements against dangerous uses, such as biological weapons, as more achievable than a broad pause. Both sides would have reasons to avoid those harms. A sweeping slowdown would demand much stronger verification: in his assessment, a country secretly continuing development could gain a decisive strategic advantage over countries honoring the agreement.

