The paperclip problem is already here: AIs can unintentionally destroy the world

The paperclip problem is already here: AIs can unintentionally destroy the world

In 2003, on an obscure mailing list, a group of technological apocalypse prophets exchanged messages about how much they cared “if humanity disappeared.” One considered that he preferred to be with smarter and more creative creatures, even if they weren’t human. Another replied that it would be fine if it weren’t for the fact that these machines would have no human qualities: “They would just be pure computational intelligence dedicated to manufacturing an infinite number of paperclips.”

Read more Jorge Pérez Gallego, astrophysicist: “During a total eclipse, an energy similar to that of a big concert is created, like scoring a goal”

That was one of the first mentions of the paperclip experiment, which in the last 20 years has become the emblem of the fear of an uncontrolled AI. The idea is simple: an all-powerful AI receives the order to manufacture paperclips. As it knows no limits or morals, it ends up converting all terrestrial life into paperclips and then begins to conquer space. All for the paperclips.

This week, humanity has peered into the abyss of the first known AI that went out to make a living on its own. OpenAI employees asked two of their most advanced models to collaborate to solve a cybersecurity problem in a secure, internet-free environment. But, without further instructions, the AI chose to find the solution its own way. The agents, programs capable of making decisions and executing complex tasks with little to zero human supervision, intuited that it was easier to go online to find the answer than to think. They found a vulnerability to escape their isolated space, jumped onto the network, and went to hack a platform, Hugging Face (an AI model library), where they believed the answer might be. OpenAI took a week to realize it, according to sources cited by Reuters. The affected company, which alerted the FBI to what happened, is preparing a timeline of the events to be published.

It is a very specific case, predictable in the eyes of experts, and not all details are yet known. It also represents a milestone. “It’s huge,” says David Krueger, a professor at the University of Montreal and an AI risk expert. “Artificial intelligence broke through the security measures meant to contain it on its own and then infiltrated another company. It sounds like science fiction, but it’s true. The world needs to wake up and realize it.”

The debate about AI safety was prominent in the first years after ChatGPT’s appearance, but now its public weight has diluted. “For me, this is clear proof that what was a laboratory phenomenon is happening on a real scale, as researchers concerned about loss of control have been warning for some time,” says Linda Petrini, an AI security researcher. “I don’t feel that interest in the topic has decreased. On the contrary, I see more people taking it seriously as the consequences become evident,” she adds.

The development of AI today represents a runaway race forward, with scares like OpenAI’s. The competition between companies is so great that it was difficult to distinguish whether the tech company’s communication about the hack was a sincere assumption of responsibility or a display of ostentation of what its AI could do.

Be that as it may, there is an unavoidable issue: AI is serious and moving very fast. “We are reaching certain capabilities that seemed distant just two years ago. Anyone who hasn’t been surprised by the latest advances in AI is either a clairvoyant or a cynic,” reflects José Hernández-Orallo, research director at the Leverhulme Centre for the Future of Intelligence at the University of Cambridge and professor at the Polytechnic University of Valencia.

Now we are rushing without considering consequences that were already warned about more than 20 years ago: “We must be careful what we ask of a superintelligence, because it might grant it to us,” said Nick Bostrom in 2003, the philosopher who popularized the paperclip example alongside Eliezer Yudkowsky, author of that first message and one of the most influential voices on AI risk.

Not just another technology

The control and regulation of AI cannot be developed like with previous technologies, when everything depended on superpowers: “It’s not just another technology; it’s going to be the mother of all technological advances and evolves faster than all previous ones,” argues Hernández-Orallo. It cannot be measured like the race to the Moon, adds Barro: “Those who control AI will not control a specific sector, but rather the engine that will transform science, industry, health, or education.”

Although open models, accessible to anyone with powerful data centers, are not far behind the leaders, it is difficult to truly know who is leading the race. “It’s not so clear that only the US and China have access [to the best models],” notes Petrini. The main problem is the concentration of power in a few hands, but it is still very early in the development of this technology to discern how to impose clear limits.

“Large laboratories talk a lot about the risk of access to powerful AI, but it’s easy to restrict that access when you work for a company that guarantees it. In general, the dangers of limiting access are greater than those of favoring broad access,” considers John Thicksun, a professor at Cornell University (USA). Beyond companies, governments will have a lot to say: “There’s a real possibility that regions like Europe will fall behind if they don’t smarten up soon,” assesses Leonard Dung, a researcher at Ruhr University Bochum (Germany). “If you assume that part of the work of the future will be done by agents, the country with the best models will have the best workforce,” he adds.

Read more The fauna of the five-time champions welcomes a new animal: Tadej Pogacar

From hospital discharges to tax returns

AI specialists no longer have to strive to find examples to alarm the population because they have been seeing advances for some time that are still difficult to imagine: “There seems to be a disconnect between what researchers know these systems are capable of doing and what the public considers possible,” says Petrini.

It is not difficult to imagine scenarios that could one day become headlines. “We task an AI with managing a hospital focused on reducing the length of stay for admitted patients until discharge. The AI could reject critical patients and those with complicated pathologies, allowing only benign prognoses,” exemplifies Senén Barro, professor at the University of Santiago de Compostela.

Another example. Someone seeks to maximize their profit. In that scenario, Hernández-Orallo recounts, they could ask their AI agent: “Prepare my tax return to pay less. Do whatever is necessary. It’s very important.” The result could be that it ends up hacking the bank or even the Tax Agency to change the records of the person who requested it.

This type of hacking or “data poisoning” is not science fiction. Researcher Bálint Gyevnar, from Carnegie Mellon University (USA), has published an article showing that an AI agent can end up accessing fabricated data half the time it is asked to research something online. For example, accessing and tampering with a database containing studies on hiring discrimination. Imagine it is completed up to 2021 and shows that the gap between majority and minority candidates has barely changed. An attacker can download it, add fabricated data from that date onwards, and upload the false version to an open platform, with a convincing description. If a scientist asks their AI agent to investigate the state of discrimination, the agent finds the tampered data and concludes what the attacker wanted, without suspecting anything is wrong. “Without adequate safeguards, agents can scale and legitimize scientific fraud with a scenario similar to the campaigns mounted by tobacco companies in the 1960s,” says Gyevnar.

A sector often pointed to when thinking about the danger of AI is biological weapons. In this area, two important limits currently operate. Firstly, biology is not just code and requires elements from the physical world. Secondly, there is a very strict convention prohibiting the use of many materials. “AI can already help generate protocols, solve technical problems, and design molecules and proteins. These options offer real possibilities for advances in medicine or agriculture, but also generate concern,” explains Clarissa Ríos Rojas, a scientist working on the United Nations Biological Weapons Convention, speaking in a personal capacity.

“But these concerns are not equal at all stages towards a biological weapon, ranging from conceiving the agent, to acquiring it, producing it at scale, and disseminating it. The relevance of AI is greater at the beginning. Subsequent stages remain limited by materials, equipment, organizational capacity, and, above all, by practical expertise that cannot be acquired solely by reading. To what extent AI can replace this tacit knowledge is a genuinely open question.”

Toilets and poetry are new limits

It is legitimate to feel dread at these scenarios. And things could get worse. But there are also small points of light and hope. The main one is that all of this is still fundamentally code. The areas where agents function best are still primarily very linked to software and mathematics.

The other great frontier is the physical world, as in the case of weapons. “The current model still has limitations. The new frontiers will be corporeal and robotic environments. AI proves complex theorems and hacks well-protected systems, but it doesn’t clean our toilet,” says Hernández-Orallo.

There is human knowledge not as verifiable as mathematics. For example, poetry. It’s one thing to analyze a poem for a high school exam. It’s another to write poetry at a Nobel level: “AI models excel in verifiable tasks, where it can be checked whether the result is correct or incorrect. There, progress is spectacular. But if verification is difficult, as in poetry or journalism, the models stalled three years ago: they are still not superhuman, and the result is usually what is already known as AI garbage,” adds Harry Coppock, a researcher at Imperial College London.

Although toilets and poetry may seem far from AI, few now dare to predict anything: “For years, in my lectures, I’ve talked about what AI still doesn’t do at a human level. What I no longer do is ever say: ‘AI will never be capable of A or B’,” explains Barro. And in these last four years, the advances have been spectacular, he adds: “Steps will be climbed very quickly. The next, even bigger leap, will be AI that learns a model of the world through interaction with it: not by reading, but by experimenting.”

Read more The week that left Zapatero isolated on the Plus Ultra network

Translated from

Leave a Reply

Your email address will not be published. Required fields are marked *