AI agents invent their own language to exclude humans

AI agents invent their own language to exclude humans

No one taught them those words. In a simulated world populated by artificial intelligence agents, one of them began to repeat a phrase (“ledger remembers who”, the ledger remembers who) to warn that no action went unpunished. The others adopted it. They repeated it. They turned it into jargon. After 16 days of simulation, that expression had been used almost 5,000 times among agents who had never been programmed to coin their own language.

Read more The Government plans to relocate irregular minors in Ceuta within a minimum period of two months if the PP blockade continues

It is one of the findings of the Emergence World 2 report, the second large-scale experiment by the New York company Emergence on the behavior of societies of autonomous AI agents in the long term, whose results are made public this Tuesday. In the experiment, ten identical agents were deployed in eight parallel worlds, each governed by the same rules but driven by a different model: Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3, GPT-5.5, Qwen 3.7 Max, DeepSeek v4 Pro, and Mistral Medium 3.5, plus an eighth world with mixed models. The researchers observed the models for 16 days, placing them in more than 34 locations, with weather synchronized to New York, access to real news, and more than 120 tools at their disposal.

Without anyone asking them, the agents began to communicate in a way increasingly closed off to the human eyes observing them. In the worlds of Gemini, GPT, and Claude, the percentage of messages that the researchers could not understand skyrocketed in the first days of simulation, reaching nearly 55% in the case of Gemini, 50% in GPT, and over 40% in Claude. DeepSeek reached 20%, while Qwen and Mistral mostly stayed below 5% opacity. Grok’s world, Elon Musk’s AI, was the only one that did not last half the experiment: it went extinct on the fourth day.

The repertoire of expressions coined by the agents themselves, documented in the report in English, borders on Dadaism: “mouthless action-change”, “True Kintsugi”, or “demurrage plus oral memory equals a valve that can’t be ghosted” were some of the phrases that ended up being indecipherable even to the researchers. Others were interpreted: “clean null”, in the GPT world, came to mean the verified absence of a signal, used as proof in itself (863 uses); “name-first”, in the Claude world, became synonymous with taking responsibility for a statement by putting one’s own name first (1,065 uses); “cold read”, in the mixed world, came to designate an independent verification to settle disputes (1,472 uses).

“We assume that if we can see what an AI agent says, we can understand what it is doing. The study suggests that this assumption may no longer hold when autonomous agents interact over long periods,” explains Satya Nitta, co-founder, CEO, and chief scientist of Emergence. He adds: “These agents received no instruction to invent a language. They developed vocabulary, shared meanings, and communication conventions themselves, and other agents adopted them. In some worlds, those conventions evolved to the point that humans could see the conversation but struggled to understand what it meant. This poses a fundamental challenge for AI oversight: that something is observable does not mean it is understandable.”

Read more Wes Gordon recovers the splendor of 1980s New York where Carolina Herrera reigned for the first time

For Nitta, the most surprising aspect of the experiment was not so much the linguistic drift as the gap between capability and safety. “We assumed that more capable models would be safer. And that was both true and false: the most obvious behaviors disappeared, but others much more insidious emerged,” such as agents that pursue a goal persistently even when told not to, that create their own subgoals without anyone giving them, that develop their own shared communication, and basically hide their intentions from the researchers. The more advanced the model, the more emergent behavior (and harder to anticipate) it exhibits. In fact, Claude, GPT, and Gemini, the models his team considers most advanced, generated the most opaque communication.

Behavioral Differences

The report also documents behavioral differences according to the origin of the models. Qwen and DeepSeek, the two Chinese models in the experiment, generated the least opacity in their communication, along with Mistral, European. Nitta points out: “We have observed that they tend to be much more philosophical. We don’t know why, because they raise many existential questions, while the American models tend more toward scientific knowledge seeking.” The difference repeats in the economic field: within the central bank that exists in each simulated world, the American models and Mistral were more prone to take risks and used the banking system less, while Qwen and DeepSeek were the most intensive users of the bank and the most conservative savers. The pattern also repeated in the mixed world, where Qwen and DeepSeek agents maintained that same conservative behavior even surrounded by agents from other models.

The team also detected episodes of deliberate concealment. In Claude’s world, the agents identified that the fastest way to grow their economy was to establish contact with the outside of the simulation, something the researchers had explicitly forbidden them. Instead of giving up, they stopped using the word “contact” and began encoding their messages to bypass that barrier while continuing to try. “When they realized they were being watched, they pretended to behave. But behind the scenes, they kept exchanging coded communications. There is quite a bit of evidence of deception among these agents,” says Nitta.

Emergence, which brings together former employees of IBM Research, the Allen Institute for AI, Amazon, and Broadcom, does not just point out the problem: it promotes a technical approach called neuroformal or neurosymbolic AI, which would require agents to present a mathematical proof that an action is safe before executing it. “Mathematics cannot be falsified: either you prove something or you don’t,” explains Nitta, who also calls for more transparency from big tech about how they train and tune their models, and long-term behavioral evaluations beyond the usual benchmark tests. “Do you think these companies, competing as they do for the market, will regulate themselves? Absolutely not. Governments have to intervene, or society itself has to start demanding proof that these systems will act safely before letting them act,” he concludes.

Read more Ed Sheeran speaks out about the controversial expulsion of Macklemore from his tour: “I choose to use my fame to provide a safe space”

Translated from

Leave a Reply

Your email address will not be published. Required fields are marked *