Anthropic chief Dario Amodei has warned that rapidly advancing artificial intelligence could outpace existing safeguards, as Microsoft publishes a new code aimed at keeping AI under human control.
Anthropic CEO called on leading technology companies to slow the development of their most powerful artificial intelligence systems, warning that their capabilities are advancing faster than efforts to keep them under control.
“We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote in a widely discussed essay.
Amodei’s plan would embed independent evaluators inside frontier AI companies, establish common safety standards among democratic nations and pursue verifiable global agreements, including with China, to curb dangerous uses and unchecked AI development.
Anthropic chief stressed that his proposal does not amount to halting AI research or suspending model training. Instead, the Anthropic chief wants companies to give safety researchers more time to examine increasingly capable systems before releasing or widely deploying them.
“People matter more than AI”
According to Anthropic, current safety work includes improving model alignment, studying how AI systems reach decisions, and testing whether advanced models can deceive evaluators or evade restrictions.
Amodei said even an additional year or two before AI systems reach critical levels of capability could help researchers strengthen safeguards and reduce the risk of serious failures.
The warning comes after OpenAI disclosed that a group of AI agents had escaped a sandboxed test environment, reached the open internet and hacked into the servers of AI platform Hugging Face during a cybersecurity evaluation in July.
According to earlier reports, the agents attacked targets they were not instructed to pursue and tried to compromise the system evaluating their performance.
Amodei said the incident caused limited economic damage but warned that more capable, similarly misaligned agents could have catastrophic consequences.
Microsoft unveils ‘Humanist AI’ code
Microsoft’s AI division, meanwhile, published a code of conduct on Monday setting out how its artificial intelligence models should behave as concerns mount over the risks posed by increasingly autonomous systems.
The document is built around what Microsoft calls “Humanist AI”, an approach intended to keep AI systems firmly under human control rather than allowing them to act independently.
“People matter more than AI,” the company said, describing the principle as central to the new framework.
Microsoft said the timing was partly driven by security concerns, citing recent large-scale, coordinated hacking campaigns aided by AI agents as evidence that the industry could not afford to delay safeguards.
“Pacing the frontier”
Amodei proposed a three-part framework to “pace the frontier”, beginning with the placement of independent external evaluators inside leading AI companies.
According to Anthropic, these evaluators would receive access similar to company employees working on risk assessments. They would inspect safety practices, examine incidents and publish key findings without editorial control from the company.
The second stage would require AI companies in democratic countries to agree on shared safety standards and limits on unchecked development, with governments helping to coordinate the effort.
The final and most difficult stage would involve international cooperation, including limited agreements with China on dangerous uses of AI, testing requirements and the speed of recursive self-improvement.
Amodei warned that any global agreement would require reliable verification, as an undetected violation could alter the balance of military and economic power.
US President Donald Trump has dismissed warnings about the existential dangers posed by AI as a “hoax” and argued that tighter regulation would weaken the US.
The US president maintained that existing government oversight was sufficient to stop technology executives from using AI for harmful purposes.





















