Anthropic's Book Scanning Draws Fresh Scrutiny

Court records show Anthropic bought millions of print books, cut off their bindings, scanned the pages, and discarded the originals for an internal research library. A June 2025 ruling found that lawfully purchased scanning workflow fair use, while July 2026 reports of unusual European antiquarian-book orders renewed preservation concerns without establishing that Anthropic was behind those purchases or that rare works were systematically targeted.
Court records show that Anthropic bought millions of print books, removed their bindings, scanned the pages, and discarded or recycled the originals while building an internal research library for AI development. The operation later became known publicly as Project Panama.
The June 23, 2025 order in *Bartz v. Anthropic* drew an important legal boundary. Judge William Alsup found the print-to-digital conversion of books Anthropic had purchased to be fair use when the physical copy was destroyed and the digital replacement was not redistributed. The same order did not excuse the company's separate acquisition of millions of books from pirate libraries.
Current reports require a separate fact check
In July 2026, German public broadcasters reported that antiquarian booksellers had received unusual bulk orders for old and out-of-print titles, often associated with Canadian reseller Zoom Books. Booksellers and trade groups raised concerns that hard-to-replace books could be entering destructive scanning pipelines for AI training.
Those reports do not establish that Anthropic placed the European orders. They also do not prove a systematic campaign to target rare books. Tagesschau said the purpose of the purchases remained unresolved after requests for comment, while WDR connected the concern to the already documented economics of destructive scanning rather than presenting evidence of Anthropic's involvement.
Zoom Books separately told Publishers Lunch in May that claims it digitizes or destroys books for AI training were false. That denial and the unresolved buyer identities are material context: the documented Anthropic operation and the current European purchasing reports should not be presented as one confirmed supply chain.
What the court actually decided
The ruling was narrower than a general license for AI companies to copy any corpus. It distinguished lawfully purchased print copies converted for internal use from pirated digital copies used to build a central library. The Washington Post's January 2026 reporting on unsealed material described how Anthropic organized the scanning project, while the court order remains the controlling record for the fair-use finding.
For data and AI teams, the practical lesson is about provenance and preservation. A defensible corpus needs records of how each source was acquired, what copies were made, who can access them, and whether the workflow destroys a scarce physical artifact. Legal permission to dispose of an owned copy does not answer the separate archival question of whether that copy should be preserved.
Key Points
- 1A federal court found Anthropic's print-to-digital conversion of lawfully purchased books fair use when the originals were destroyed and the digital replacements were not redistributed.
- 2July 2026 reports documented unusual European antiquarian-book orders but did not establish that Anthropic placed them or that rare works were systematically targeted.
- 3Training-data governance should track acquisition, copying, access, and preservation separately because a copyright ruling does not resolve cultural-archive risk.
Scoring Rationale
The documented scanning workflow and court ruling are relevant to AI corpus provenance, copyright controls, and physical-source preservation. The current European purchasing reports add timely scrutiny, but their buyer identities and connection to Anthropic remain unproven, which limits the event's certainty and impact.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems


