Some AI engineers are afraid of what they’re building. They want to speak up while they still have leverage
Former Anthropic researcher Jacob Coxon has opened the floodgates.
Former Anthropic researcher Jacob Coxon has opened the floodgates. Coxon resigned last week in a viral thread on X that accused AI companies of not acting responsibly, even though they all believe AI could trigger the end of humanity. We’ve heard predictions like this before from across Silicon Valley, but Coxon’s posts triggered a groundswell of pressure on AI companies and governments to do something about the pace of development and safety. And much of that pressure is coming from staffers inside the companies as it is from outside of them. One staffer at a top AI company told CNN the fears of how AI could hurt humanity keeps them up at night. Another researcher who recently left a different AI company said it’s a common subject of conversation at parties and social events in Silicon Valley. “You can’t spend more than a few hours in this community without realizing that a very substantial number of people are really pretty worried about these sorts of outcomes,” said the researcher. “A majority would say there’s some chance of it killing everyone.” Rising worries Coxon’s posts have been viewed more than 170 million times and broke through to the general public in a way that previous warnings about AI capabilities did not. That may be because they came just weeks after stunning revelations of AI agents going rogue, escaping their test environments and hacking into other systems completely autonomously. Dozens of Coxon’s colleagues in the AI industry publicly supported his statements, with some making even more dire predictions. The groundswell of concern seemingly forced leaders of AI companies to react – Anthropic CEO Dario Amodei sat down for interviews over the weekend, and OpenAI CEO Sam Altman and SpaceXAI CEO Elon Musk agreed to Amodei’s proposal to embed independent watchdogs at the AI companies. Current and former AI staffers told CNN in recent days that they’re worried about the extremely fast pace of AI development. While the rest of the world is awed and alarmed by AI models solving nearly 100-year-old math problems and swarms of AI agents conspiring together to hack into another company’s systems, these researchers fear for the day when the AI systems can improve themselves without human input, which is known as recursive self-improvement. An AI system that can improve itself creates an automatic feedback loop that could not only bring the AI to superhuman capabilities, but do so in such a short window that it would be difficult or too advanced for humans to monitor and intervene if necessary. “For the first time, I am asking myself if things are moving too fast. I’m honestly not sure, but I am sure that it would be good for us to have an answer to ‘What would a successful pace look like?’” OpenAI VP of Research Aidan Clark posted on X last week. ‘I would burn my equity to the ground’ For many AI scientists, there is a dilemma: They can stay at the top AI companies, where they can make gobs of money and try to reduce risks from the inside. Or they can leave (sometimes in protest) and try to make a change from the outside, such as at one of several organizations that conduct outside research and evaluations of the AI models. “I would burn my equity to the ground in a heartbeat for a 1% higher chance we make it out of this situation alive. I expect a great many of my colleagues across the industry would as well,” wrote Drake Thomas, who works on AI safety at Anthropic. “I promise you, we are actually just f**king scared, it’s not galaxy brained marketing.” Amodei expressed no issue with Coxon’s resignation post, telling CNN’s Anderson Cooper this week that he agrees “with Jacob much more than I disagree with him.” For years, AI researchers and engineers have been a premium, scarce resource that have demanded high salaries and high leverage, which gives them the freedom to speak out. But AI staffers told CNN that they fear that as AI gets better at training and improving itself, their leverage goes down, prompting today’s urgency. “As we get into this recursive self-improvement loop, I think that might substantially reduce staff’s bargaining power, because frankly, you’ll be able to replace many of the staff with models that can do as good a job,” the researcher who recently left a top AI company said. The former researcher compared the atmosphere inside leading AI labs to the Manhattan Project, the World War II effort to build the first atomic bomb, where scientists pursued a technology of extraordinary power while fearing its destructive potential. But these staffers justify working on such potentially destructive technology because “it will be better if we develop it than if our adversary develops it – and therefore we have at least moral permission, if not a moral obligation, to do it.” Overplaying the risks But not everyone in the AI industry agrees with these concerns. Some said that much of the doom conversations are rhetorical and that there are a lot of “selection effects” based on where the people work and their social circles. Anthropic, for example was created by a group of former OpenAI staffers that wanted to prioritize safety precautions and ethical guardrails. Current and former staffers at Anthropic say existential risks caused by AI are a “a constant topic of conversation.” That’s not the case at every company, however. Since Coxon’s posts went viral, many of the posts echoing his views have come from staffers at Anthropic, OpenAI and Google. But fewer have seemingly come from Musk-led xAI or Mark Zuckerberg’s -Meta. Many in the industry agree that a lack of AI safety standards could cause real-world harm, like unintentional hacking of critical infrastructure. But some are skeptical of the current doomsday outlook and think the reality is more nuanced. “Sorry, but asking Jacob (Coxon) about AI extinction risk is like asking your AC guy about climate change,” wrote Hugging Face CEO Clement Delangue, whose company servers were hacked by rogue OpenAI agents this year. “Not saying it’s necessarily uninteresting or wrong per se but let’s keep things in perspective and hear from the full range of expertise across the ecosystem!”
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.