Anthropic spent millions buying and physically destroying millions of physical books. The books were not read; they were shredded after being scanned. The intended reader was not a human, but a machine learning model. This is not a dystopian fiction. This is the latest evolution in the AI data acquisition race.
As a protocol PM who once spent weeks auditing multi-sig contracts in Frankfurt, I've seen technical choices mask ethical ones. But the destruction of physical books for training data is a different kind of vulnerability—one that targets not a smart contract, but the very fabric of cultural memory. The legal shield is a 2025 court ruling that allows 'one-to-one replacement' of physical books with digital copies for non-distributive purposes. On paper, it is fair use. In practice, it is a license to incinerate provenance.
Context: The Decentralization Philosophy of Data
Blockchain and AI share a foundational belief: information should be free, transparent, and verifiable. But this story exposes a contradiction. The decentralized ethos demands trustless verification, yet the most trusted data source for AI training is now a physical object—a book—that must be destroyed to become digital. The court's logic is simple: if you own a physical book, you can convert it to a digital format, provided you destroy the original to keep the copy count consistent. This is the same reasoning that allows libraries to digitize rare manuscripts. But when applied at scale by corporations with billions in capital, it becomes a system of cultural extraction.
The service enabling this is ISBNdb, which markets itself as a one-stop shop for 'destructive scanning.' They buy books, cut off the binding, run high-speed scanners, and then shred the remaining paper. The remains are verifiably destroyed, often with a notarized certificate. The selling point? Clean data. Books published before 2022 are free from AI-generated text and modern data poisoning techniques. For an industry increasingly worried about model collapse from synthetic data, this is gold.
Core: The Technology and Values Analysis
Based on my experience auditing data pipelines for decentralized protocols, I see three layers to this trend. First, the technical layer: destructive scanning bypasses copyright by exploiting a legal loophole. The digital copy replaces the physical one in a one-to-one ratio. But in practice, a digital file can be copied infinitely, and the court's ruling does not police the subsequent use of that file for training large language models. This is not a technical limitation; it is a governance gap.
Second, the values layer: AI companies are prioritizing data purity over cultural preservation. They argue that using physical books ensures the training data is authentic human expression, free from the noise of the internet. But this authenticity comes at the cost of erasing the physical artifacts that carry it. A first edition with marginalia, a signed copy by a marginalized author, a book from a small press that never digitized—all are sacrificed at the altar of algorithmic alignment.
Third, the economic layer: this creates a new asset class—'training-ready text'—that is finite and non-renewable. Once a book is shredded, its content can never be verified again against its physical source. The market for such data is opaque, with ISBNdb not disclosing pricing. But the scarcity of clean, pre-AI books will drive up costs, making it a privileged resource only for the well-funded.
I recall a conversation with an artist on the Art Blocks platform in 2021. She said, 'The code is the art, but the soul is the artist's intent.' Here, the code is the training data, and the soul is the cultural context. By destroying the physical book, we lose the provenance of its existence—who owned it, where it was printed, how it survived. Code has conscience.
Contrarian: The Pragmatic Test
One might argue that this is an overreaction. AI companies need clean data to build safe models. Books are being digitized anyway by Google and others. Shredding the original is a legal necessity to avoid copyright infringement. Isn't this just a more radical version of scanning?
The counter-point is that Google Books scans and preserves the digital copy without mandatory destruction. Libraries keep both. The destructive model destroys the physical evidence of human creation, and the digital copy is not treated as a cultural artifact but as a commodity—a token to be consumed. There is no backup plan for the metadata of the physical item: its binding, its marginalia, its ownership history. Everything that makes a book a unique object is lost.
Moreover, the assumption that pre-2022 books are 'clean' is itself a blind spot. Books contain biases, errors, and outdated worldviews. Training on them without cross-referencing digital sources may produce models that are historically accurate but socially regressive. The search for purity becomes a path to rigidity.
Takeaway: The Vision Forward
The real question is not whether AI should use books, but how we preserve cultural provenance in a digital age. Blockchain offers a path: tokenizing the digital copy with irrevocable metadata about the physical source, including its destruction? That is a contradiction. Better to create a public registry of destroyed books, with hash-linked digital copies stored in decentralized storage networks like IPFS or Arweave. This would allow verification without destruction. Liquidity flows where belief resides.
If AI companies cannot trust the internet's data, they should not trust the destruction of books either. The solution is collaborative: digitize with consent, share the digital copies under open licenses, and keep the physical books in libraries. That is the only way to train ethical AI without burning our cultural inheritance. Trust is the new token.
The court ruling may be legal, but it is not ethical. As builders of decentralized systems, we must ask: what are we preserving, and what are we losing? The answer will define the authenticity of our AI for decades.