Somewhere in Europe, an old book leaves a shelf. It might be a technical manual from the 1970s, a forgotten novel, a regional history, or a scientific monograph nobody has requested in years. The order does not look like the work of a collector.
Antiquarian booksellers across Europe have reported unusual purchasing patterns: long lists of unrelated titles, obscure books, automated-looking orders, buyers apparently more interested in ISBNs than editions, and transactions whose shipping costs sometimes seem disproportionate to the value of the books themselves. Suspicions have grown that at least some of this activity is connected to the artificial-intelligence industry's search for training material.
Those suspicions should be handled carefully. The ultimate destination of individual books is often impossible to establish, and companies associated with unusual bulk purchasing have denied acquiring books for destructive AI digitisation.
The underlying phenomenon, however, is no longer hypothetical. We know that an AI company has already bought physical books on an industrial scale, cut them apart, scanned them and converted them into machine-readable data. Anthropic called the operation Project Panama. The objective was not to build a great library for humans; it was to build one for machines.
For years, the AI debate has concentrated on models and compute. The next strategic resource may be much older than either of them: human knowledge itself.
The easy corpus has already been collected
The first generation of large language models benefited from a historical accident. Humanity had spent roughly three decades putting enormous quantities of human-created material online — websites, forums, Wikipedia, newspapers, research papers, government documents, books, software repositories, blogs, reviews, manuals, conversations. Nobody built the internet as an AI training set. That is what part of it became: an enormous corpus of human language that could be collected at machine scale.
The largest models consumed astonishing quantities of this material, and then an obvious problem appeared. There is only so much high-quality human-created text in existence. Worse, the internet itself is changing. Generative systems now produce articles, product descriptions, comments, summaries, marketing copy, code, translations, images and increasingly entire websites. The corpus is beginning to contain the output of the models trained on the corpus.
Researchers have been studying the consequences. Training future models recursively on model-generated material can degrade the underlying distribution of knowledge: rare patterns disappear, and outputs become increasingly concentrated around what previous models already considered probable. Under some training regimes, researchers describe the resulting degradation as model collapse.
The details are more nuanced than the headline. Synthetic data can be extremely useful when it is generated, verified and mixed carefully with authentic data. But the signal holds: original human data does not become less valuable as AI improves. It may become more valuable, because the world is producing less and less of something specific — data created by humans before generative machines participated in its creation.
What an obscure book from 1983 is now worth
Consider a book published in 1983: a history of hydraulic engineering in Switzerland, perhaps, or a collection of interviews with textile workers, or an obscure study of Alpine agriculture. Perhaps only 1,500 copies were ever printed. For decades its economic value was negligible; a second-hand bookseller might price it at CHF 12.
To an artificial-intelligence system, that same book has several unusual properties. It is almost certainly human-authored. Its provenance can be established and its publication date is known. Its language has not been contaminated by generative AI. It may contain specialist knowledge poorly represented online, regional terminology, unusual sentence structures, forgotten technical vocabulary or perspectives that never migrated to the web. And there may be no digital copy anywhere.
The dusty shelf is no longer merely storing a book. It is storing authenticated human data, and that changes the economics. A book can be nearly worthless to the conventional information economy while being useful to the intelligence economy.
This helps explain one of the stranger episodes revealed by recent AI litigation. Anthropic did not merely download books; it bought them, millions of them. For Project Panama, physical books were acquired in enormous quantities, their bindings removed so the pages could be scanned efficiently, and the resulting text incorporated into a digital research library. The physical object was expendable. The information inside it was not.
The value migrated from the object to the corpus.
Buy, scan, digitise, train
There was also a legal dimension. In Bartz v. Anthropic, a US federal court considered whether Anthropic's use of copyrighted books for training constituted fair use, and the ruling drew distinctions worth keeping intact. The court found the use of books for training Anthropic's models to be transformative fair use in the circumstances before it. It treated Anthropic's conversion of lawfully purchased physical books into digital copies for an internal research library differently from its acquisition and retention of millions of books from pirate libraries.
Those distinctions rule out the slogan version of the case — that AI companies are simply allowed to steal books. That is not what the decision established. But the economic incentive the legal landscape creates is worth stating plainly: for certain uses, physically purchasing a book and converting it into training data can provide stronger provenance than downloading the same text from an unauthorised online collection.
The industrial process therefore becomes buy, scan, digitise, train. The second-hand book market has become part of the AI supply chain. That would have sounded absurd five years ago.
The scramble for the undigitised world
Reports from antiquarian booksellers read differently in this context. Booksellers have described unusual orders for obscure and apparently unrelated titles; one company repeatedly discussed in reporting is the Canadian bookseller Zoom Books. Zoom Books has denied that it purchases and destroys books for AI training, and there is insufficient public evidence to attribute every unusual purchasing pattern to an AI company.
When this volume was published, the destination of individual bulk antiquarian-book purchases was often impossible to establish. On 17 August 2026, 404 Media reported that a bookseller placed a tracker inside an anonymous order of roughly 1,000 books and followed the shipment to Amazon's VGT3 facility in Las Vegas. Workers interviewed by the publication described books being received, having their bindings removed, and being scanned. Amazon confirmed that it purchases books through commercial channels to develop and improve products and services. The reporting concerns commercially acquired old, scarce and out-of-print books — not unique manuscripts or irreplaceable cultural artefacts.
The destination is no longer entirely hypothetical. A physical-book-to-machine-corpus supply chain can now be observed directly.
The larger economic logic does not depend on solving the mystery of any individual buyer. AI developers want high-quality human-created corpora. Much of the easily accessible digital corpus has already been collected. Copyright litigation has made provenance more important. The open web is becoming progressively contaminated with synthetic material. And enormous quantities of human knowledge still exist only in physical or poorly digitised form.
The consequence is a new search frontier. The machine has read much of the internet; now it is looking for everything the internet forgot — books, microfilm, broadcast recordings, regional newspapers, specialist journals, government archives, university collections, oral histories, maps, photographs, audio, video, scientific collections, private archives.
Cultural institutions that once appeared peripheral to the technology industry may find that they are sitting on some of its most valuable raw material.
Switzerland has been recording itself for decades
For decades, SRG SSR has recorded this country: Federal Councillors, referendums, elections, wars, economic crises, football matches, farmers, scientists, artists, protests, weather reports, debates, interviews, regional events and ordinary people. Ordinary people speaking — in German, French, Italian, Romansh and Swiss German, with regional accents and dialects that are dramatically underrepresented in the global digital corpus.
We have understood that archive primarily as media history and cultural heritage. Artificial intelligence gives it another meaning: it is also a dataset. Not a generic one. It is a longitudinal, professionally produced, culturally specific record of Switzerland observing and describing itself.
In August 2026, SRG Director General Susanne Wille revealed that SRG intends to make its archive available for the development of Apertus, the Swiss public large language model built by researchers associated with ETH Zurich and EPFL. Her argument was that the archive was financed by the Swiss public and should therefore be used within a transparent Swiss environment. At approximately the same time, SRG announced its intention to restrict uncontrolled crawling of its online content by AI companies.
The two positions only look contradictory. SRG is not saying that nobody may use its knowledge. It is saying that it should have agency over the conditions under which its knowledge is used. That is a sovereignty question.
Openness without agency
Digital sovereignty is frequently confused with isolation: build everything domestically, store everything domestically, ban foreign providers, close the borders around data. That is a poor definition. A small country such as Switzerland cannot and should not attempt technological autarky. It will use foreign chips, software, research, models, clouds and services. Interdependence is not the opposite of sovereignty. Uncontrollable dependency is.
The same principle applies to knowledge. A sovereign knowledge resource does not need to be secret, or even restricted. It may be openly accessible, licensed internationally, used for academic research, used by commercial AI systems, or freely downloadable. What matters is whether the institution responsible for it retains the ability to determine the terms on which it is made available.
The difference is between deliberate access and uncontrolled extraction. Openness without agency becomes extraction. Protection without access becomes stagnation. The workable policy space lies between them.
The knowledge extraction economy
Consider the value chain. Swiss taxpayers fund universities, broadcasters, libraries, archives and scientific institutions. Swiss researchers produce knowledge, Swiss journalists document society, Swiss authors write books, Swiss institutions preserve historical records, and Swiss citizens generate language, culture and debate. A foreign technology company collects those resources and converts them into training data. The training data improves a proprietary model. The model becomes a commercial service. Swiss companies, governments and citizens then pay to access that service.
Swiss society creates knowledge
↓
external platform extracts knowledge
↓
platform converts knowledge into machine capability
↓
Switzerland rents the capability back
There is nothing inherently wrong with international companies learning from Swiss knowledge. Knowledge has always travelled; science, culture and the internet all depend on it. But the structure should look familiar. Countries have encountered versions of it before: raw materials are exported, higher-value products are manufactured elsewhere, and the finished products are imported. The difference is that the raw material is no longer iron ore, oil or timber. It is accumulated human knowledge.
For two centuries, industrial policy asked who controlled coal, steel, electricity, oil, manufacturing capacity and transport infrastructure. The intelligence economy introduces another strategic resource: the corpus.
From cultural heritage to productive infrastructure
This requires us to reconsider what an archive is for. A national library preserves books because societies believe their intellectual history should survive. A broadcaster preserves recordings because they document public life. A federal archive preserves government records because institutional memory matters. A university preserves research because knowledge should remain available to later generations. Those purposes remain. AI adds another: these collections can become productive infrastructure, because they can teach machines. That makes their quality, provenance, accessibility and governance economically significant.
Switzerland already holds substantial reserves. The Swiss National Library collects books, newspapers, magazines, maps, official publications, websites and images. The Swiss Federal Archives preserve the documentary history of the Confederation. Universities maintain research repositories and specialised collections. Libraries preserve regional publications, cantons and municipalities maintain their own archives, and museums hold cultural collections. SRG holds decades of audiovisual material; the Swiss National Sound Archives preserve recorded sound; parliamentary records preserve political discourse; scientific institutions hold specialist datasets that may have no equivalent anywhere else.
Individually, these are archives. Collectively, they are part of Switzerland's knowledge infrastructure.
National knowledge reserves
Switzerland understands strategic reserves. We maintain systems designed around the possibility that critical resources — food, fuel, medicines, energy — may become unavailable. The objective is not self-sufficiency. It is resilience.
Artificial intelligence suggests a related concept: national knowledge reserves. Not a giant government database, not a ministry of truth, not a compulsory national training corpus, and certainly not a patriotic model trained to produce officially Swiss answers. A national knowledge reserve would be a governance concept. It would identify culturally, scientifically, linguistically and institutionally significant corpora and ensure that Switzerland retains durable, machine-readable access to them.
Such reserves would need different access classes. Some material could be completely open, some available for research, some licensed, some restricted by copyright, some subject to privacy safeguards, and some should never become AI training material at all. Sovereignty does not mean indiscriminate ingestion. It means retaining the ability to decide.
The corpus needs an exit test too
Infrastructure sovereignty has an obvious test: if your provider disappeared tomorrow, could you continue operating? The same question belongs to AI knowledge infrastructure.
Suppose Switzerland comes to depend heavily on foreign foundation models. That may be perfectly reasonable. But suppose those models become unavailable, prohibitively expensive, geopolitically restricted or commercially unsuitable. Could Swiss institutions train, adapt or retrieve against alternatives? That depends on whether we hold the underlying corpora, whether they are machine-readable, whether their provenance is known, whether their licences are documented, whether they may legally be used, whether they can move between model architectures, and whether the metadata needed to interpret them exists. Could another provider use them? Could an open model? Could a future Swiss model?
If the answer is no, we may possess archives without possessing AI-ready knowledge sovereignty. Digitisation alone is not sufficient. The next problem is portability of knowledge into machine systems.
The provenance premium
In an internet increasingly populated by generated material, provenance itself becomes valuable. A newspaper archive from 1994, a parliamentary transcript from 1972, a radio interview from 1961, a scientific journal printed in 1987, a book deposited with a national library — for each of these we know approximately who produced it, when, in what institutional context, and often through what editorial process.
Compare that with a random webpage discovered in 2031. Was it written by a human, generated by a model, or generated by a model trained on another model's output? Automatically translated? SEO spam? Copied, manipulated, reconstructed from something else?
The problem is not only whether information is true. It is whether its lineage is knowable. That gives archives another increasingly scarce property: trusted provenance. The AI economy may therefore place a premium not on data volume but on data lineage. The future may belong less to whoever holds the most tokens than to whoever holds the most trustworthy ones.
Why the long tail matters
This also explains why obscure books matter. Large language models are exceptionally good at representing the centre of a distribution: the common language, the popular concepts, the frequently repeated knowledge. But cultures are not made only from their most common expressions. They also live in the tails — regional dialects, minority languages, forgotten terminology, local history, technical specialisms, unfashionable ideas, small communities, rare professions, historical disagreements, things that were published once and never repeated online.
When recursively generated data begins amplifying what models already consider probable, those tails matter more, not less. A corpus containing the obscure and the unusual is not merely larger; it is more diverse.
That makes libraries unusually well suited to this moment. They were never designed to preserve only what was popular. They preserve what existed. That is precisely the behaviour an intelligence infrastructure needs.
Who teaches the machine about Switzerland?
Imagine a child born in Basel in 2035. When she wants to understand the 1992 EEA referendum, she will probably not search a newspaper archive; she will ask a machine. When she wants to hear how people in the Bernese Oberland spoke fifty years earlier, she will ask a machine. When she wants to know why Switzerland built nuclear power stations, how women's suffrage developed, what happened during the Jura conflict, how Swiss industry changed, or what political arguments surrounded immigration in 2014, she will ask a machine.
The model will answer. Where its understanding came from matters more than whether the model itself is Swiss. A model developed in California but grounded transparently in high-quality Swiss sources may provide better Swiss knowledge than a model developed in Zurich without them. Sovereignty cannot be reduced to the flag attached to the model developer.
Who teaches the machine about Switzerland?
And who controls the memory from which it learns?
Not a national truth machine
Once governments begin discussing "national knowledge", the concept can become uncomfortable quickly. A democratic knowledge infrastructure must not become an official version of reality.
Switzerland's archives contain disagreement, and that is a feature: left and right, urban and rural, federal and cantonal, German-speaking and French-speaking, majorities and minorities, winners and losers, correct predictions and spectacular mistakes. A sovereign corpus should preserve those contradictions.
The objective is not to teach machines what Switzerland thinks. There is no such thing. It is to preserve enough evidence that machines can understand what Switzerland has thought. One of those ambitions produces propaganda. The other preserves memory.
Apertus is not the point
It would be easy to read SRG's announcement as a story about Apertus. Apertus may succeed. It may fail. It may become an important European public model, or be surpassed by something better. Models are replaceable: architectures change, parameter counts grow, training techniques improve, and today's frontier becomes tomorrow's commodity.
The archive is not replaceable in the same way. You cannot recreate seventy years of Switzerland. You cannot regenerate interviews with people who have died, reproduce historical broadcasts after the fact, manufacture authentic Romansh radio from 1963, or buy another copy of a conversation that happened once.
Models are reproducible in principle. History is not.
That makes the archive potentially more durable, strategically, than the model consuming it. Apertus is not the most interesting asset in this story. The corpus is.
The strategic inversion
For much of the current AI boom, we have assumed the hierarchy of the stack looks like this:
chips → compute → models → applications → data
The hunt for books and archives suggests another reading. Chips can be manufactured again. Compute capacity can be expanded. Models can be retrained, software rewritten, applications replaced. Authentic historical human data cannot be re-created. Every conversation that was never recorded is gone. Every discarded local newspaper, every book destroyed without digitisation, every dialect recording that was never preserved may be gone. Some information resources are effectively non-renewable.
That produces an inversion. The oldest component of the AI stack may turn out to be the hardest one to reproduce.
The last human corpus
Humanity is not going to stop producing authentic culture. People will continue writing books, doing research, making films, having conversations and creating art. But the informational environment has changed permanently: human and machine production will increasingly intermingle, and the clean historical separation is gone.
Everything produced before generative AI therefore belongs to a distinct period — not because it was better, or because humans are intrinsically more truthful, but because we can know something specific about its origin. It came before the machines began writing back.
That gives the historical corpus a new technological property. It is not merely old. It is pre-generative. Seen that way, the scramble for old books becomes easier to understand, and so does the value of newspaper archives, radio, libraries and SRG's tapes.
We spent decades thinking about these institutions as places where human memory was stored for humans. Another reader has arrived. It reads faster than any human could, it can absorb millions of documents, and what it reads will influence what billions of people subsequently hear.
What Switzerland should actually do
Switzerland does not need to close its archives. It should do close to the opposite: recognise their value, digitise aggressively, preserve originals where appropriate, improve metadata, document provenance, resolve rights, develop machine-readable access, create transparent licensing models, support research, enable open models, allow responsible commercial use, protect personal data, preserve minority languages, and prevent irreversible exclusive capture.
And above all, retain options.
The question is not whether Swiss knowledge should train artificial intelligence. It already will. The question is whether Switzerland participates deliberately in determining how. SRG's decision is therefore more interesting than a partnership with a Swiss AI model: it may be an early recognition that cultural archives have become part of technological sovereignty.
Other institutions should pay attention, because somewhere else another old book is leaving a shelf. Someone has worked out that the forgotten parts of human culture have acquired a new value. The intelligence economy is looking for raw material, and much of it has been sitting quietly in our libraries all along.
Conclusion — Memory is infrastructure
We spent thirty years digitising culture because we wanted machines to help humans find knowledge. We may spend the next thirty discovering that we digitised culture so machines could learn it.
That does not have to be frightening. Machines able to access centuries of human knowledge could make obscure research discoverable, help preserve minority languages, reconnect fragmented archives, translate forgotten texts, and give ordinary people access to collections previously available only to specialists. Those possibilities make governance more important, not less.
The strategic AI debate therefore needs to expand beyond GPUs, datacentres, cloud sovereignty, foundation models and inference. There is another layer underneath all of them: memory. Countries that understand this early will not necessarily build the largest models. They may hold something more durable — the ability to decide how their accumulated knowledge participates in the intelligence economy.
Because the strategic question is no longer only who owns the model.
Who owns the memory from which the model learns?
And who gets to decide what happens to it?
Sources
-
404 Media, 17 August 2026 — We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility Primary investigation: a tracker placed in an anonymous order of roughly 1,000 books arrived at Amazon's VGT3 facility in Las Vegas; workers described books being unbound and scanned. Amazon confirmed that it purchases books through commercial channels to develop and improve products and services. 404 Media
-
Forbes, 17 August 2026 — AI Companies Are Buying and Destroying Antique Books. Here's Why Secondary coverage of the same 404 Media investigation, including the tracked shipment and Project Panama context. Forbes
