Anthropic's Book Scanning Draws Fresh Scrutiny

Anthropic bought and destructively scanned millions of print books for AI training, a workflow described in June 2025 court reporting and found fair use by Judge William Alsup under specific conditions. A 2026 Revolver News article broadened the discussion to claims about rare-book destruction, but the supplied material does not independently establish an industry-wide campaign targeting rare books.
Anthropic bought and destructively scanned millions of print books for AI training, according to court documents reported by Ars Technica in June 2025. A Revolver News article later renewed attention to the practice, framing it as a threat to rare and irreplaceable works. The strongest source material supplied concerns Anthropic's book-digitization operation rather than a newly documented industry-wide campaign involving rare books.
Ars Technica reported that Anthropic purchased millions of print books, cut their bindings, scanned the pages, and discarded the physical copies to create training material for its AI systems. According to Ars Technica's account of the court record, Anthropic hired former Google Books partnerships head Tom Turvey in February 2024 and tasked him with obtaining "all the books in the world."
The legal finding
Ars Technica reported that US District Judge William Alsup found Anthropic's destructive scanning workflow to be fair use under the specific facts before the court: Anthropic had legally bought the books, destroyed each physical copy after scanning, and retained the digital files internally rather than distributing them. Ars also noted that the ruling did not absolve alleged earlier piracy involving other training data.
The legal distinction matters for dataset builders. A court finding about format conversion from lawfully acquired copies does not establish a blanket rule for acquiring, reproducing, distributing, or training on every type of copyrighted corpus. Provenance, access method, retention, and downstream use can each change the legal analysis.
Claims about current buying activity
Two 2026 commentary articles describe purported current bulk purchases of older physical books. Futuro Prossimo attributes claims of orders ranging from 1,000 to 1 million books to ISBNdb, and identifies Anthropic's disclosed scanning operation as a "pilot case." A Medium article reports that antiquarian booksellers in Germany, Switzerland, and Austria have discussed large orders for used non-fiction, citing estimates of roughly 700,000 volumes in Germany and as many as three million worldwide.
Those volume estimates and the identity of buyers are not independently substantiated in the supplied material. Neither the Ars Technica reporting nor the cited court finding establishes that rare books were being targeted, nor does it document a current industry-wide purchasing program.
For ML practitioners, the episode illustrates a continuing data-governance problem: high-quality, pre-generative-AI text is valuable for corpus curation, but physical-source workflows add chain-of-custody, licensing, digitization-quality, and cultural-preservation questions. Comparable digitization programs often separate the need for machine-readable text from the preservation of physical artifacts through non-destructive scanning or archival partnerships.
Key Points
- 1Court records described Anthropic's destructive scanning of millions of legally purchased books, creating a concrete precedent for training-data provenance discussions.
- 2Judge Alsup's fair-use finding was tied to internal retention and purchased copies, not a general authorization for all AI training corpora.
- 3Comparable digitization efforts often distinguish machine-readable text acquisition from preservation of physical artifacts through non-destructive archival workflows.
Scoring Rationale
The underlying court disclosure is relevant to training-data sourcing, copyright compliance, and corpus provenance. However, the newest claims about widespread rare-book destruction are not independently supported by the supplied reporting, and the principal legal event occurred in 2025.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

