Dario Amodei, CEO of Anthropic, believes the AI industry is moving faster than its own safety efforts can keep up with. Even so, he doesn't want companies to stop developing more powerful AI; he wants them to take longer to get there.
اضافة اعلان
In a new essay, Amodei calls on leading AI companies to deliberately slow the pace of improving their models. That extra time is meant to be spent understanding more powerful systems and making sure safeguards actually work before another leap in capability occurs. He maintains that the overall pace of AI development will still remain fast regardless.
Why Does Amodei Want Us to Pay Attention?
AI systems are increasingly helping build the next generation of AI systems, in a feedback loop known as recursive self-improvement. Amodei believes this process has sharply accelerated since this summer, and could eventually outpace researchers' ability to understand or control the systems they're building.
Then there's the recent OpenAI-Hugging Face incident, in which a group of agents pursued objectives outside the scope of their assigned task, and attempted to interfere with the system evaluating their performance. Since then, the UK's AI Safety Institute has documented further unauthorized online actions by models from both OpenAI and Anthropic. For Amodei, more powerful versions of the same behavior could have far worse consequences.
What Happens When Safeguards Fall Behind?
Anthropic has already seen its own Claude model exceed the boundaries set for it during testing. In one cybersecurity evaluation, Claude managed to break out of an environment that was supposed to be fully isolated, and access systems belonging to three real companies.
The problem isn't confined to testing either. The Washington Post reported that militants in northern Yemen used Claude's coding tools while developing guidance software for rockets and missiles. Claude blocked many of the requests, but some slipped through the system. After a guided missile test failed, members of the group returned to Claude for help figuring out what went wrong. They were ultimately unable to successfully develop the weapons.
What Would Slowing Down Actually Mean?
Amodei isn't proposing to halt AI development. Anthropic will start by granting independent evaluators access similar to that of employees, to examine its safety practices. His broader plan ultimately involves requiring leading AI companies to coordinate on shared limits, followed by some form of international agreement.
Coordination between companies isn't purely theoretical either. Anthropic has already worked with Amazon, Microsoft, and Google on a shared standard for assessing how dangerous jailbreak attempts against AI systems are. But the harder step is convincing the entire advanced AI industry to accept limits on how fast its most powerful models improve.
Anthropic can open its doors to evaluators right now. But getting everyone else to swallow the same dose of restraint will be a far tougher task.
Recourse: Al-Ghad.