Why Anthropic is destroying books — Kathryn James traced it
Kathryn James follows Project Panama from warehouse cutters to court, explaining why clean human prose became premium AI material.
Anthropic bought perfectly usable books, sliced off their spines, scanned the loose pages and threw away the remains. Vandalism usually involves less procurement paperwork.
This operation had a codename: Project Panama.
An internal memo surfaced through Bartz v. Anthropic and examined by Kathryn James in The Guardian on August 5, 2026. Its mission statement was admirably unhinged:
Project Panama is our effort to destructively scan all the books in the world.
The memo also explained the codename:
Why use a codename? … [B]ecause we don’t want it to be known that we are working on this.
I’ve spent 20 years building and shipping technology. When a company creates a codename, hires logistics people, rents warehouse space, compares scanning vendors and develops a legal theory, I call that a strategy.
So, why is Anthropic destroying books? Kathryn James followed the paper trail to a boring answer with fairly terrifying consequences: scanner economics and American copyright doctrine pointed toward the same hydraulic paper cutter.
Project Panama was an industrial operation
According to James’s review of the court records, Anthropic hired an experienced logistics manager, sourced books through vendors, brought in staff and stored the inventory in a warehouse. Specialist contractors handled the digitization.
Court exhibits showed organized shelves packed with labelled books. Employees moved between stacks and carts while volumes waited for processing. Somebody planned intake volumes. Somebody calculated scanning throughput and disposal capacity.
I’ve dealt with physical workflows while building products such as ALYT, where hub hardware had to cooperate with a mobile app and cloud infrastructure. Every warehouse process hides a thousand annoying decisions. Someone approves the equipment. Someone signs the vendor contract. Someone eventually discovers that the cheap scanner jams every 300 pages and has ruined everybody’s Tuesday.
Judge William Alsup described the vendors’ work:
stripped the books from their bindings, cut their pages to size, and scanned the books into digital form – discarding the paper originals
Loose pages fly through high-speed document scanners. Bound volumes need slower equipment, usually overhead cameras, flatbed scanners or V-shaped systems that support the binding while each page is photographed.
Project Panama chose speed.
I understand the calculation. A startup buying millions of books will model the cost per scan, failure rate, staffing bill and warehouse time. Preservation becomes an expensive spreadsheet row unless an executive explicitly protects it.
The secrecy is harder to excuse. Anthropic felt confident enough to spend heavily on the operation while recognizing that footage of pallets entering a spine cutter would play badly outside the building.
And the memo did say all the books in the world. I grew up in Ivrea, the town of Olivetti, so I have a weakness for grand engineering ambitions. Even by Olivetti standards, this needed an espresso and a second read.
Companies do not accidentally build book-destruction supply chains. Procurement teams compare bids. Lawyers review exposure. Executives decide that negotiating with authors creates more friction than buying used copies and cutting them apart.
That final decision is where things get ugly.
AI slop made old books more valuable
Anthropic wanted books because they are excellent training material. They contain edited prose, sustained arguments, unusual language and stories that survive beyond six seconds of TikTok attention.
The court materials described what Anthropic valued:
well-curated facts, well-organized analyses, and captivating fictional narratives
Anthropic hoped this material would help Claude write with the accuracy and appeal of human authors. Apparently the AI industry rediscovered books after spending a decade training everybody to communicate through push notifications, Slack reactions and videos of strangers reorganizing refrigerators.
Bravissimo.
The pre-2022 cutoff matters. A physical book published before modern generative AI exploded is unlikely to contain paragraphs produced by ChatGPT, Claude or Gemini. Paper has become a rough certificate of human origin.
That certificate now carries a premium because the open web is filling with synthetic text. Models generate articles, product descriptions and forum answers. Later models scrape the same material and consume it as training data. Repeated often enough, that cycle can contribute to model collapse, with systems learning from increasingly distorted synthetic distributions.
TechRadar’s July 2026 analysis connected demand for pre-2022 books directly to this risk. Tom’s Hardware also reported that professionally edited print sources offer cleaner human-authored material than large parts of the current web.
The feedback loop is magnificently stupid. AI companies helped flood online spaces with cheap generated prose. That pollution made old human writing strategically scarce, so the industry returned to secondhand bookstores to extract it.
Molto Silicon Valley.
My first reaction was dismissive. I assumed people were getting sentimental about paper, like my refusal to throw away a stained Marcella Hazan cookbook when the same recipe exists online.
I was wrong.
Access to clean human language now affects who can build the strongest models. Anthropic’s private corpus can improve Claude while remaining unavailable to researchers, libraries and smaller European AI companies. The physical copy leaves circulation. Its commercial value stays inside one American corporation.
Everyone who declared long-form writing obsolete should sit with that for a minute. Some of the richest technology companies on Earth are spending millions to ingest books because sustained human thought is premium raw material.
Copyright law rewarded the cutter
Judge Alsup concluded that Anthropic could purchase a physical book, convert it into a digital copy for internal model development and destroy the paper volume. Under the facts before him, the private scan replaced the purchased copy.
He summarized the logic in four words:
One replaced the other.
Alsup also emphasized that Anthropic kept the collection closed:
There is no evidence that the new, digital copy was shown, shared, or sold outside the company.
Destroying the book helped Anthropic’s legal argument. Keeping the intact volume alongside its scan would leave two usable copies. Discarding the paper allowed the company to characterize its workflow as format conversion: one purchased object entered, one private digital file survived.
American law does not require every scanned book to be destroyed. Alsup’s district court ruling evaluated Anthropic’s specific process, and it is not a binding nationwide appellate precedent. Another judge can reach another conclusion.
Corporate incentives rarely wait for the Supreme Court to publish a tidy FAQ, though. Lawyers already have a favorable pathway to study.
Buying one used book, making one internal scan and eliminating the original can strengthen a fair-use defense while cutting scanning costs. Copyright law has given the hydraulic cutter legal value.
Google Books offers a useful comparison. Google’s library-scanning project generally borrowed books, photographed their pages with non-destructive camera systems and returned the volumes. Tom Turvey, who previously worked on Google Books partnerships, later became involved with Project Panama, according to court reporting summarized by GIGAZINE and Tom’s Hardware.
Anthropic bought used books, cut them apart and kept the resulting corpus private. The warehouse received waste paper. Anthropic received proprietary training data.
I’ve built enough products to know what happens after a court validates the cheaper workflow. Finance teams turn it into a spreadsheet template. Vendors package it as a service. Competitors quietly copy it while their communications departments practice saying “digitization initiative” with a straight face.

Project Panama converted a purchased physical book into a private digital file. The scan survived. The book did not.
Image alt text: Why Anthropic is destroying books through Project Panama’s destructive scanning process.
The $1.5 billion settlement taught one lesson
The settlement and the destructive-scanning ruling cover separate pools of books.
Anthropic had accumulated more than 7 million pirated books. Authors Andrea Bartz, Charles Graeber and Kirk Wallace Johnson sued the company in 2024, arguing that their work had been obtained from pirate libraries and used for Claude AI training data.
The court treated those files differently from books Anthropic purchased and scanned. Model training and the conversion of legally bought books received favorable fair-use treatment. Downloading and retaining pirated copies did not.
Anthropic agreed to a $1.5 billion settlement in 2025. The deal received final approval on July 27, 2026, covering approximately 500,000 works at roughly $3,000 per eligible title. PC Gamer reported that claims had been filed for 91% of eligible works.
That is a huge payment. The lesson for every other AI lab still fits on a Post-it: obtain a receipt first.
Kathryn James reports that Anthropic CEO Dario Amodei viewed copyright clearance as a drawn-out legal and business slog. I recognize the founder frustration. Negotiating rights across thousands of publishers, estates and territories sounds like spending eternity trapped in an airport lounge with DocuSign.
Creators live on the other side of that “slog.” Their books supply the language that improves a commercial model, while a used-book seller receives the only mandatory payment.
The settlement compensates claims tied to piracy. It does not create a general requirement for AI companies to negotiate with an author before training on a lawfully purchased copy.
Creative Bloq made the distinction sharply in its analysis: the case punished how Anthropic assembled part of its library. It changed far less about what Anthropic could do after legally acquiring a book.
Because the dispute settled, no appellate court produced a rule binding every future AI copyright case. Google, Meta and OpenAI still face their own disputes, each with different facts.
Authors won money. AI labs received a procurement lesson.
Strange buyers are clearing old shelves
Project Panama is documented. A second trend is murkier: booksellers in Australia and Europe have received large, seemingly random orders for obscure titles.
Guardian Australia reported on August 2, 2026, that Delfina Manor, who runs Good Reading Secondhand Books in Benalla, Victoria, received a Zoom Books order for around 30 to 40 books packed into three boxes.
Manor described the operational reality:
They ordered, paid in advance, and they didn’t quibble over the postage … [but] it would have been about 30 to 40 books, and so finding them, packing them, making sure you hadn’t missed one out … It was just driving me nuts,
John Sainsbury of Sainsburys Books in Melbourne reported similarly random, price-insensitive demand for niche and decades-old stock. One order combined a 1970s soil-mechanics manual, Born to Thunder: Champions of New Zealand Cycling, a 1982 collection of Early Australian Poetry and a local history of Hawthorn.
I would love to meet the human reader with that exact weekend planned.
The poetry collection sold for $9 after sitting on a shelf for about 20 years. Here the economics become emotionally messy. A bookseller finally moves dead inventory, while nobody knows whether the volume will reach a reader, an arbitrage warehouse or an industrial cutter.
Nick and Jenny Dawes estimate that Grant’s Bookshop holds around 500,000 books across its warehouses. Nick told The Guardian he would feel conflicted if somebody bought the entire stock to cut it up. He would refuse a sale if the buyer planned to destroy the only known copy.
Fair enough. Used-book stores need revenue, and a 1987 engineering manual does not pay rent by developing a distinguished layer of dust.
The uncertainty creates the preservation risk. Anthropic says it has never purchased from Canadian reseller Zoom Books and that its acquisition programs do not buy or destroy rare or antiquarian titles. Zoom Books says it resells books intact and does not digitize them. The company declined to identify customers because its commercial agreements are confidential.
I cannot responsibly connect those Australian orders to Anthropic. The evidence does not support that claim.
ISBNdb adds another odd detail. Archived marketing reported by 404 Media and PC Gamer advertised printed-book sourcing for LLM training, with orders ranging from 1,000 books to 1 million. The page promised customer confidentiality:
your identity, strategy, and acquisition targets are never disclosed.
Its archived copy understood the optics perfectly:
‘AI company destroys two million books’ is not a headline that generates sympathy.
ISBNdb later removed the offering and said it had merely tested market interest. The company stated that it had never purchased, scanned or sold a physical book.
Meanwhile, an anonymous specialist bookseller told 404 Media that weekly sales had jumped from around 20 books to hundreds after the purchasing surge began. The orders shared little beyond the presence of ISBNs, which suggests lists assembled from bibliographic databases.
No Library of Alexandria cosplay is required here. An opaque market can quietly remove uncommon books while sellers remain unable to judge the preservation risk. Confidential acquisition targets make basic stewardship almost impossible.
A corporate corpus makes a lousy library
Anthropic has described plans for a “forever” research library. Bold. My self-hosted Docker stack occasionally develops opinions after a routine update, so I treat eternity claims from private technology companies with some caution.
A preservation library documents where its material came from and protects exceptional copies. It also creates a credible path for future access. A corporate training corpus guarantees none of this.
The public does not have a complete inventory of the titles Anthropic destroyed. Researchers cannot freely inspect the scans. Future historians cannot examine the original physical copies.
OCR captures text. It misses plenty.
Books can contain marginal notes, ownership stamps, printing errors and handwritten recipes. Paper can reveal a production method. A binding can show whether an edition circulated cheaply or targeted wealthy buyers. An inscription might connect the object to a family, an institution or a political movement.
I think about cookbooks because I am Italian and therefore legally obligated to turn every technology argument into lunch. A stained community cookbook with somebody’s handwritten substitution for lard tells me how a family actually cooked. Clean OCR gives me the official recipe and deletes the improvisation.
Tim White, who runs Melbourne’s Books for Cooks, collects material showing where food and society meet. He told Guardian Australia that rare booksellers handle objects with stories.
His verdict on one-off volumes was simple:
If it’s a book that has that storytelling element to it, or it’s a one-off, it’s horrific.
Rare bookseller Gwenyth Todd of Chatelaine Books screens buyers because she finds it physically abhorrent when books are cut apart and their illustration plates sold separately. Profitable book destruction predates AI. Language-model companies can industrialize it at a scale plate dealers could only dream about.
The United States has relatively few protections governing books as cultural heritage, especially compared with safeguards for other artifacts. There is no practical endangered-species list for an obscure regional history with three known copies.
Any AI lab buying physical books at industrial scale should follow six rules:
- Scan scarce, annotated, antiquarian and out-of-print works without cutting their bindings.
- Check library catalogues and bibliographic databases for rarity before destruction.
- Publish an inventory of every title destructively scanned.
- Deposit preservation-grade files with a trusted library under controlled access where copyright requires it.
- Offer authors and publishers direct licensing routes.
- Allow independent audits of acquisition vendors and scanning facilities.
These requirements cost money. Good. A company building a billion-dollar model from humanity’s written record can afford a rarity check.
On July 27, 2026, Elon Musk said he had instructed the SpaceXAI team to preserve rare books and scan them “the hard way.” I’ll accept that as a useful minimum commitment. Patron saint of librarians would be a little premature.
The equipment already exists. Google used non-destructive scanning. The Internet Archive uses preservation-oriented systems. Datamation Information Services, one of the scanning companies connected to Project Panama through court reporting, offers destructive and non-destructive methods.
Choosing the cutter is a business decision.
The most advanced language machines on Earth are rummaging through used bookstores because they still cannot manufacture what they need most: a deep record of humans thinking before the machines arrived.
We spent years hearing that books were dead, writers were replaceable and everything meaningful could fit inside a feed. Now pre-2022 human writing is valuable enough to warehouse, scan and feed into products worth billions.
Before any AI company destroys a book, I want two answers. Is this copy replaceable? Who gets access to what survives?
By 2028, I expect major AI procurement contracts to include rarity screening and public title inventories, either through regulation or publishers refusing to cooperate without them. If they don’t, our descendants may discover that we preserved the written record of humanity inside proprietary model weights while feeding the readable copies into a recycling bin.
Frequently asked questions
Why is Anthropic destroying books?
Anthropic destroyed purchased books because removing their bindings enabled faster, cheaper high-speed scanning. Destroying each paper copy also supported its argument that one purchased object had been converted into one private digital file rather than duplicated, strengthening the fair-use position accepted by Judge William Alsup.
Did the Anthropic copyright settlement cover books it legally purchased and scanned?
The $1.5 billion settlement covered claims involving pirated books, not the separate pool of physical books Anthropic legally purchased and scanned. The court treated downloading and retaining pirated files differently from converting purchased books into private digital copies for internal model development.
Why are pre-2022 books valuable for AI training?
Pre-2022 books offer edited, sustained human writing that is unlikely to contain text generated by modern systems such as ChatGPT, Claude or Gemini. As synthetic prose spreads across the web, older printed books provide cleaner human-authored material and reduce exposure to recursive training on AI-generated content.
Sources
- Why is Anthropic destroying books? | Kathryn James
- ‘More than just objects’: Australian booksellers raise alarm over ‘horrific’ destruction of rare titles to feed AI
- Company Offering Printed Books to Train AI Stops After 404 Media Coverage
- AI companies are anonymously buying and destroying millions of books through middleman services to avoid headlines about AI companies buying and destroying millions of books
- Company that said it could scan and destroy books for AI data-harvesting has deleted that part of its website: 'no such service was ever brought to life'
- AI companies are reportedly shredding millions of books after using them to train AI models — tech giants outsource to middlemen to secretly buy up books for training material