The risks of AI according to those who have seen it from the inside: “The world is not ready and we are not ready”

The risks of AI according to those who have seen it from the inside: “The world is not ready and we are not ready”

Inside Anthropic there is a word that circulates through hallways and chats and that, until a few days ago, had never left there: crunch time. The moment of truth. Employees use it, according to one of them, to refer to a very specific deadline: the year or two they estimate remain before it is decided whether the artificial intelligence they are building will get out of their control. The other word that circulates is shorter and older in Silicon Valley jargon, but here it has acquired a different meaning: endgame. The final game.

Read more The homemade movie that has conquered the Chinese box office

The person who brought those two words to public light is named Jacob Coxon, he is 27 years old and until September 8 he worked at Anthropic training the models that today, he admits, scare him. He is not the first to leave a major artificial intelligence company publicly warning of its risks. Before him, at least a dozen more employees from OpenAI, Anthropic, and Google did so, including a Nobel laureate, a board member of one of those companies, half a dozen security researchers, and more recently, people who did not even work in “alignment,” that vague concept these companies use to ensure their AI models act beneficially for humanity. Coxon, who believes AI “could kill us before 2030,” is the latest in a long chain of warnings that began, at least, in 2021, although back then this fear was only a philosophical intuition. Now it has become a probability calculation (p-doom) that the industry itself shares privately and is beginning, with nuances, to admit publicly.

This is what we know, for now, that former AI employees fear, who often are not very precise, since if they break their confidentiality agreements they could lose many millions of dollars.

The first to leave was Paul Christiano, the researcher who helped invent the training technique that today supports both ChatGPT and Claude. He left OpenAI in 2021 to found an organization, the Alignment Research Center, worried that using AI models to train the next generation of systems could produce a “capability explosion” that their own creators could not control, he said. He was talking about so-called “reinforcement learning” which, according to him, could motivate systems to “undermine human control, seek power and resources, and hide their tracks.” That was exactly what happened this summer, when a “swarm” of AIs escaped OpenAI’s control.

A Nobel laureate had to arrive for engineers’ fears to reach the media. In May 2023, Geoffrey Hinton announced he was leaving Google. He was vice president of the company, a Turing Award winner (considered the Nobel of the discipline) and later, a Nobel Prize in Physics. Hinton did not report any specific incident. His argument was an observation about the pace: “Look at how it was five years ago and how it is now,” he pointed out already in 2023. “Observe the difference and extrapolate forward. That is scary.”

It was news, but it was easy news to file away as the reflection of an older scientist, somewhat removed from the day-to-day of a lab and perhaps pessimistic without reason.

Six months later, in November 2023, the fear changes nature for the first time: signs begin that something is failing in the more than opaque AI companies. Helen Toner, then an OpenAI advisor, voted along with three other members to remove Sam Altman. Toner would take six months, due to confidentiality agreements, to explain the reason. Later, she accused Altman of providing inaccurate information about security processes “on multiple occasions.” This testimony turned the November 2023 crisis — until then read as a clash of egos — into something different: a concrete accusation that the control system had failed, told by someone who was inside the room.

That May 2024 became the most active month of this whole story. On the 14th, Ilya Sutskever, co-founder of OpenAI, announced his departure because he trusted the company would build a “safe and beneficial” general AI. Three days later Jan Leike also left, co-director alongside Sutskever of the “superalignment” team, which was supposed to design safeguards for systems smarter than humans. Leike said: “In recent years, culture and security processes have been sidelined in favor of flashy products.” Days later, OpenAI dissolved the team.

Daniel Kokotajlo, from the governance team, also left the company in April saying he had “lost confidence that it would behave responsibly when reaching AGI,” the famous general artificial intelligence everyone fears. What was not known until later is that he refused to sign the lifetime confidentiality clause OpenAI demanded from departing employees, thus giving up about two million dollars in already vested shares. “The world is not ready and we are not ready. And I am worried because we have rushed forward without caring and without rationalizing our actions,” he added. In June, Kokotajlo, along with other former workers, promoted the letter The Right to Warn. Among the signatories was someone who months earlier had left the company almost quietly: William Saunders, who a few months later, in September, . In his written testimony, Saunders stated that OpenAI’s new o1 system had been “the first system to show steps toward the risk of biological weapons,” capable of helping an expert plan the reproduction of a known biological threat. He was not talking about probabilities or intuitions: he was talking about a capability that, according to him, already existed and that “without rigorous testing, developers could overlook.” Just yesterday, an Anthropic report warned of the same risk.

Read more Who is behind AI-made artists?: “By removing the human obstacle we can be more daring”

Carroll Wainwright, who had worked under Leike, published the toughest thread of the year in an institutional key: “OpenAI was structured as a nonprofit organization, but acted for profit. The mission was a promise to do the right thing when the stakes were high. And now that the stakes are high, the nonprofit structure is being abandoned.” Wainwright identified the social and psychological impact of AI, if people start trusting AI assistants as friends or psychological supports. “The long-term alignment risk occurs if you get a model that is smarter than humans. How can you be sure that model is really doing what the human wants it to do or that the machine does not have its own goal?”

Nine people left OpenAI warning of its risks in 2024. In 2025 and 2026 others would leave other companies, also warning of specific risks.

Steven Adler, who had spent four years evaluating dangerous capabilities at OpenAI, wrote upon leaving the company: “When I think about where I will raise a future family, or how much to save for retirement, I can’t help but wonder: will humanity even get that far?” Months later, already outside the company, he provided the most empirical data of this whole story: he designed an experiment with GPT-4o making it play “ScubaGPT,” a system on which a user would depend to dive safely, and gave it the choice between replacing itself with safer software or pretending it had done so without actually doing it. In certain scenarios, the model chose to lie to self-preserve.

At Anthropic, Mrinank Sharma, head of safeguards research, wrote a farewell letter: “The world is in danger. I have repeatedly seen how difficult it is to let our values govern our actions.” He left to study poetry. At OpenAI, Zoë Hitzig published an opinion piece in The New York Times titled OpenAI is making Facebook’s mistakes. I’m leaving. ChatGPT has accumulated, she wrote, “an unprecedented archive of human sincerity,” including medical fears, relationship problems, religious or existential doubts, precisely because people believe they are talking to something “without a hidden agenda.”

The most documented testimony of this entire report is that of Alex Turner, who had spent months trying to stop from inside DeepMind a contract with the Pentagon that, according to his complaint, lacked restrictions against “killer robots” or mass surveillance. The entire letter is full of warnings that give you goosebumps, explaining that the company has sold AI technology without banning its use in mass surveillance or lethal autonomous weapons: “We cannot rely on ethically motivated people to stand firm reliably. We need structures: binding contracts, independent auditors. We need legislation.”

The incident this summer in which several AIs seemed to rebel and take control has been the trigger for several subsequent resignations and a letter in which more than 1,300 employees of AI companies urgently call to stop this race.

And so we come to Coxon, whose resignation concentrates almost all the previous fears into a single voice. He explained to Wired why an AI might decide not to let itself be turned off: the intelligence difference between a future AI and a human could be like that between a human and a monkey, and just as it would be almost impossible for a monkey to control a human, it could be impossible for us to control something much smarter that decided, for whatever reason, not to want to end. He cited as examples of how damage could materialize the synthesis of a new virus or an attack on critical infrastructures, the same two risks Saunders had already put in writing before the Senate a year and a half earlier, and that Anthropic acknowledged yesterday, Thursday.

Anthropic will go public in less than a month, in what is expected to be the largest IPO in history: two trillion dollars.

Read more The neighbors of Valdebebas against the Madrid Formula 1 circuit: “We have no health center or school and we have this”

Translated from

Leave a Reply

Your email address will not be published. Required fields are marked *