Publishers Seek Sanctions Against OpenAI Over Chat Logs

News publishers led by The New York Times asked a Manhattan federal judge on July 9 to sanction OpenAI, alleging it concealed internal copyright-search tools and deleted or made unsearchable billions of relevant ChatGPT conversations. The motion cites a roughly 78-million-conversation dataset and Project Giraffe; OpenAI rejects the claims as false and says the dispute threatens user privacy.
News publishers led by The New York Times asked a federal judge in Manhattan to sanction OpenAI on July 9, alleging discovery misconduct in their copyright litigation. The latest motion is an accusation from the plaintiffs, not a court finding, and OpenAI disputes it.
What the publishers allege
The publishers say OpenAI represented that searching its training data and ChatGPT output logs for their copyrighted material was infeasible or unduly burdensome. They argue that later testimony showed the company had already conducted internal searches and maintained tools for measuring when models reproduced source material.
TechCrunch reports that the motion cites a database of roughly 78 million de-identified ChatGPT conversations and an internal effort called Project Giraffe. The publishers allege Project Giraffe included a Bloom-filter-based system for detecting and recording regurgitated material. They also allege OpenAI deleted or made unsearchable billions of conversations that could have been relevant to discovery.
The requested remedies include legal fees and limits on how OpenAI may use a previously produced 20-million-log sample. The court has not ruled on the latest request.
OpenAI's response and the court record
OpenAI called the accusations false and said the publishers are continuing efforts to obtain private user conversations as their case weakens. That response matters because the dispute combines two separate questions: whether relevant evidence was preserved and produced, and how courts should protect user privacy when ordering discovery from AI systems.
A March 9 court order, available through Justia, had already directed OpenAI to produce reservoirs of 78 million and 10 million logs under a de-identification protocol. The same order directed further production and discussion concerning Project Giraffe. That earlier order confirms the discovery dispute, but it does not decide the publishers' new sanctions allegations.
Why this matters for AI teams
The practical lesson is about evidence governance rather than the underlying copyright merits. Model-evaluation datasets, output logs, filters, deletion jobs and internal measurement tools can become litigation records. Teams operating generative-AI systems need documented retention rules, reproducible searches and clear ownership of legal holds. Those controls help distinguish ordinary privacy-driven deletion from conduct that could later be challenged as evidence loss.
Key Points
- 1Publishers asked a federal judge on July 9 to sanction OpenAI for alleged discovery misconduct; the court has not ruled on the request.
- 2The motion cites a roughly 78-million-conversation dataset and Project Giraffe, while alleging billions of other conversations were deleted or made unsearchable.
- 3OpenAI rejects the allegations and says the publishers' demands threaten the privacy of users who are not parties to the lawsuit.
Scoring Rationale
The sanctions motion concerns evidence preservation and internal AI-evaluation systems in a major copyright case. Its industry impact is meaningful for model governance and retention practices, but the allegations remain disputed and the court has not ruled.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems
