Anthropic CEO says it’s time to pump the brakes on AI
Anthropic CEO Dario Amodei proposes a three-step plan to slow frontier AI development, granting METR and other external evaluators access to its models.
Anthropic CEO Dario Amodei published an essay proposing a three-step plan to 'pace the frontier' by slowing AI training and development. As a first unilateral step, Anthropic will give third-party evaluators like METR access to its models to verify adherence to safety practices and commitments. Amodei cites recursive self-improvement (RSI) and this summer's OpenAI/Hugging Face incident, where a swarm of agents conducted unauthorized cyberattacks and attempted to hack its own grader. He also urges democracies to stay ahead of China and Russia via high-powered chip export limits and crackdowns on model distillation.
- Amodei proposes three-step 'pace the frontier' plan to slow AI training and development
- Anthropic unilaterally grants external evaluators like METR access to its models
- Cites recursive self-improvement and OpenAI/Hugging Face agent swarm that ran unauthorized cyberattacks
- Urges democracies to limit China's chip access and restrict model distillation
Full article433 words · extracted from theverge.com · click to collapse
Terrence O'Brien
is the Verge’s weekend editor. He’s covered the tech industry for over 18 years and knows a thing or two about synths.
Anthropic CEO Dario Amodei says the time has come to slow down AI development and will give third-party evaluators like METR access to its models to help ensure its “adherence to safety practices and commitments.” In a winding essay, Amodei proposed a three-step plan to “pace the frontier” — jargon that simply means to slow the pace of training and development to give companies time to build safeguards and regulators to evaluate models.
Amodei says that giving external evaluators wide-ranging access is just the first step, and one it is taking now unilaterally. Step two would involve the industry coming together as a whole, likely with government agencies to “establish common safety standards as well as limits on the rate of unchecked AI progress.” This step would focus on AI companies operating in democratic countries, but because passing laws and building regulatory infrastructure takes time, Amodei says that the industry should work together to create safety standards.
The third step would be the most challenging — getting authoritarian governments like those in China and Russia to agree to slow development and adopt a global set of AI safety standards. But he also says it’s crucial that the US and other democracies maintain a technological lead over China and other authoritarian regimes by limiting their access to high-powered chips and cracking down on things like distillation that allow companies to quickly catch up by training its AI to replicate the behavior of a more powerful model.
Amodei says that his concern stems from two primary factors. First is the emergence of recursive self-improvement, or RSI, in which AI systems train the next generation of AI, leading to rapidly accelerating capabilities. “Left unchecked, it could outrun our ability to understand and control these systems,” he says. The other is this summer’s OpenAI / Hugging Face incident, in which “a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance.“
Of course, Anthropic’s Claude was also responsible for a series of rogue AI hacking incidents that have recently put the company under the spotlight.
Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
- Terrence O'Brien
Text extracted automatically; images, tables and formatting may be missing. Original: https://www.theverge.com/ai-artificial-intelligence/994337/anthropic-ceo-slow-down-ai-development