When internet knowledge is no longer enough: AI sets its sights on second-hand bookstores

When internet knowledge is no longer enough: AI sets its sights on second-hand bookstores

Marçal Font was surprised one day to detect some strange movements in his Fènix bookstore in Badalona. “They bought four books for 35 euros with shipping costs of 67 euros. One minute later they ordered another book with 20 euros more in shipping costs. And so on. It made no sense,” he recalls. It was then that he contacted his friend Xavi Vinaixa, a researcher and technical director of a technology company, to find out what could be behind it and they soon found the answer to the mystery: a company was placing orders to digitize the books and incorporate them into artificial intelligence systems.

Read more Spencer Tunick takes his collective nudes to the Canary Islands: «There is too much hatred towards open-minded people»

Through an international forum, they got in touch with European booksellers who were experiencing the same thing. Purchases were made through ZoomBooks (a Canadian company for buying and selling second-hand books) and shipments were sent to intermediary logistics operators such as PrepFort, in Illinois, which offers jobs to scan, catalog, and record data accurately and quickly. Both the booksellers and the researchers consulted maintain that this structure makes it difficult to identify the final recipient of the copies, although everything points to technology companies linked to the development of artificial intelligence.

The suspicions of the booksellers found resonance in an investigation by The Washington Post about Anthropic, the company that created the Claude model. Based on court documents related to copyright lawsuits, the newspaper revealed the existence of an internal program known as Project Panama, described by the company itself as an effort to digitize printed books on a large scale in order to incorporate them into their artificial intelligence systems. According to the documentation analyzed by the American newspaper, the process is based on the destructive scanning of physical copies.

Destroying books to obtain data

The image inevitably recalls Fahrenheit 451, although here the books are not destroyed to censor ideas, but to extract data. “Unbinding a book by guillotining the spine to free its pages allows the copy to be scanned and digitized faster and cheaper than the conservative and careful scanning used by libraries and archives,” points out Vinaixa, although it implies the disappearance of the physical copy. It is paradoxical that a project created to feed on data ends up destroying data along the way.

Unlike Bradbury’s novel, here no one wants to eliminate the content. On the contrary, what they want is to absorb it. “For years, AIs were trained with huge amounts of text available on the web. Now many companies face several problems because they have already consumed much of the quality public text available and need more diverse, specialized, and better-written data to try to avoid homogenization, cognitive biases, errors, and the informal vocabulary that predominates on the internet,” adds Vinaixa.

That is why books are now so valuable to technology companies: they contain more elaborate, edited, and structured knowledge than much of the content on the internet. “They buy specialized books, with low commercial turnover and great documentary value: local studies, technical publications, traditional gastronomy, or essays that often only interest researchers,” points out Font. They are now buying books with ISBNs, which guarantees that much of the books acquired are part of a relatively well-cataloged ecosystem protected by legal deposit systems.

But what will happen if companies want to continue expanding their repositories? Where does the demand of technology companies for data end and when does the content created by humanity begin to be protected? “This is my biggest concern. They will end up reaching much less protected and identified materials: fanzines, local publications, neighborhood bulletins, or documentation of social movements that never entered conventional bibliographic circuits,” adds Font.

Moreover, there are increasing copyright litigations. In December 2023, The New York Times sued OpenAI and Microsoft for using its content to train generative artificial intelligence systems. The newspaper claims that millions of copyrighted articles have been used to train chatbots that now compete with it. In April 2024, eight U.S. newspapers joined the lawsuit.

Read more FIFA already has the reports to investigate Argentina’s behavior in the World Cup final and assess possible sanctions

The great collapse

If AIs end up absorbing much of the available human corpus, one possibility is that they will increasingly train on content generated by other AIs. This phenomenon already has a name in research: model collapse. The hypothesis is that if one generation of models learns from the production of the previous generation, knowledge may progressively degrade. It is a kind of photocopy of a photocopy of a photocopy. “That is why out-of-print books represent a reserve of human content not yet contaminated by artificial intelligence,” points out Antonio Ortiz, analyst and communicator specialized in artificial intelligence and technological trends.

“Now, if it is true that lack of data is a problem within the sector, it is also true that it is not as central as it was two years ago,” he adds. “The most advanced models are improving thanks to reinforcement learning. Previously, the industry sought more data volume and greater computation, whereas now it bets on specialized training based on quality data and well-solved problems. Data remain important, but the race is shifting towards increasingly qualified data,” he adds.

The privatization of knowledge

The great scientific advances of the 20th century happened thanks to data, universities, and public research, which was also preserved in public institutions. But something seems to be changing: there is beginning to be a shift of discovery, development, and scientific preservation towards a certain privatization.

“The main risk is a privatization of knowledge. The logical thing would be to first use human content to train large public language models, like the European Mistral AI, because knowledge must continue to be a public good, responding to logics of transparency, authorship recognition, and sustainability,” argues Patrícia Ventura, PhD in Communication, Media, and Culture from the Universitat Autònoma de Barcelona.

“All the magnificent work being done to preserve heritage today, AI does infinitely faster,” says Font. “The question is not whether AI is destroying heritage, but whether public institutions will be able to identify, catalog, and preserve heritage before machines turn it into private data,” he adds. For this, it is important to understand what is at stake at this moment.

And it is not the dystopia posed by Bradbury—they do not want to burn books to make them disappear—but rather to absorb their data and privatize knowledge. Because, unlike the controversial project started by Google twenty years ago, Google Books, this time they do not intend to give us back access to the digitized primary source, but to allow us access to transformed knowledge: that which the machine regurgitates after swallowing the book.

Read more Emptied Hispania: Roman remains of small towns become engines of rural development

Translated from

Leave a Reply

Your email address will not be published. Required fields are marked *