OpenAI, the company behind ChatGPT, announced last week that one of its models had been able to solve the so-called “existence and smoothness problem” of Navier-Stokes, one of the seven Millennium Problems of the Clay Institute, considered the most complex in the history of mathematics. And there are rumors that they are about to unlock another. The Navier-Stokes problem had been waiting to be solved for 90 years. The company led by Sam Altman says it achieved this in 88 hours using 10,000 artificial intelligence (AI) agents, although it has already acknowledged that it applied the method developed by two Spanish scientists and still has not clarified whether it also used the work of another researcher from Anthropic, its main competitor. “We did not use their prompts [questions or requests] nor their tests to guide our models or agents,” said an OpenAI spokesperson regarding the latter. And added that, “although unlikely, we cannot rule out that data derived from the use of our products helped improve our models.”
Read more The road to prison for Trompas, the confessed killer of Claudia Tacoronte
This last point is key. OpenAI does not make it clear whether the use that a scientist — or any other user — makes of its models becomes part of the tool, that is, if it is incorporated into the model. “OpenAI is very opaque about what it does with user interactions,” points out Julio Gonzalo, professor of Languages and Computer Systems at UNED and deputy vice-rector for research.
Could AI chatbots then steal an idea from one user and offer it to another? There is no simple answer to this question. First, because we know less and less about how these models work. And second, because the companies that develop these tools are quite opaque about the data sources they use to feed their models.
Model training
Before an AI model can be used, it must undergo the so-called training phase. This process consists of applying an algorithm to a gigantic database to perform a series of actions. Deep learning, which is one of the most used AI strategies today, applies neural networks (several layers of interconnected models) directly on a data set to extract patterns. The system learns autonomously, without needing to be programmed, although the results can later be manually refined (supervised learning) so that the system sees which are good and which are not, and persists along the indicated path.
For these systems to work, three things are needed: a large computing capacity (to carry out the training process and inference, that is, managing each request or prompt against the finished model), the technical formulation of good algorithms, and databases large and relevant enough so that what the models learn makes sense.
The latter is crucial. The better the data worked with, the better the model will be. And the problem is that data is starting to run out. According to some estimates, the GPT4 model (which is two versions behind the most powerful ChatGPT) was already fed with all the content on the internet, as well as many other files. AI developing companies are looking for new ways to feed their databases. “It is not entirely known what OpenAI or Anthropic do for their models to learn. We only have intuitions and what those who have worked there say. It is quite clear that they train on conversations people have with them, but theoretically not on all,” says Álvaro Barbero, director of the AI Lab at the fraud prevention company Lynx Tech and professor at Afi Global Education.
“User interactions are a very valuable source of information for training models,” notes Gonzalo. “That information is integrated with the immense amount of data the models have, which mainly comes from the internet, along with some that may be generated by the company’s own workers,” points out Carlos Gómez Rodríguez, professor of Computing and Artificial Intelligence at the University of La Coruña and expert in natural language processing, the branch of AI that seeks to understand and generate texts.
Read more The United States sent military engineers this week to inspect an old base in southern Greenland
Could it be the case that a researcher writes prompts containing valuable ideas, and those ideas appear to another user after a few months, when a new version of the model updated with those requests is available?
“On paper, it could be that if the prompts used to train the model contain useful information to solve a particular problem, it is applied in subsequent requests. It is not something that can be demonstrated 100%, but it is logical to think so,” Gómez indicates. But there is also no certainty that all prompts are used. “Today, there is no way to know if any particular prompt was represented in the model. That trace cannot be followed, because during the training process the model’s weights are modified, which are like a huge matrix of numbers that we can’t even interpret,” Gómez adds.
“The learning process performs a kind of lossy compression, and the rarer or less frequent what it reads is, the more likely it is not to memorize it well,” explains Gonzalo. That is: the rarer prompts have less chance of being incorporated into the model. “But the opposite also happens: the more specific and rare the new user’s conversation is, the more likely the model is to remember conversations about that rare and specific topic. Overall, it is possible, though not inevitable, that the original ideas you share with a model may end up being used by others,” he concludes.
Data control
If what users type can be used by the model to train it and later regurgitate it to another, should companies fear for the security of their trade secrets?
OpenAI’s or Anthropic’s solution lies in payment. “If you use the free version, they give you no guarantee that they won’t use your data to train, meaning they probably are,” Barbero illustrates. “In theory, if you are a company and sign a professional agreement with OpenAI or Anthropic, they are supposed to have clauses that prevent them from using your conversations to improve their models. I say in theory because, for example, Anthropic had to pay compensation to book authors when it was discovered they had downloaded many from pirate sites.”
Returning to the resolution of major mathematical problems, Gómez emphasizes that AI had already solved others before Navier-Stokes. He also adds that experts on that particular problem say the model also made contributions (not just followed the research line of others). It would also not be unusual for special attention to be paid to contributions in the form of prompts from the most capable users. “OpenAI is offering prominent scientists, mathematicians, and engineers free access to its cutting-edge models,” Gonzalo points out. “Thinking they do it altruistically would be very naive, with everything we are seeing.”
Read more These professionals earn more than a minister and almost no one knows it