Major news publishers accused OpenAI of retaining roughly 78 million ChatGPT conversation logs in a de-identified database, even after the company had represented that the data was gone. The allegation surfaced in a sanctions motion filed in The New York Times Co. v. Microsoft Corp., et al., No. 1:23-cv-11195, a copyright case pending in the U.S. District Court for the Southern District of New York. The filing, led by The New York Times and joined by other publishers, asks the court to penalize OpenAI for what the plaintiffs describe as discovery failures tied to data the company said it had deleted.
Why 78 million retained chat logs change the copyright fight
The sanctions motion puts a specific number on a problem that has shadowed OpenAI’s legal defense for months. According to court records filed in the Southern District, OpenAI maintained a reservoir of roughly 78 million ChatGPT conversation logs in de-identified form. The publishers argue that OpenAI told opposing counsel and the court that these records no longer existed, only for the database to surface during discovery.
That gap between what OpenAI said and what the filing shows matters for two reasons. First, if the court agrees the company misrepresented the status of these logs, sanctions could range from monetary penalties to adverse inference instructions that tilt future jury deliberations against OpenAI. Second, the sheer volume of retained conversations raises pointed questions about how the company handles deletion requests from different groups of users. If 78 million logs survived after OpenAI indicated they were removed, the obvious follow-up is whether individual users who asked for their data to be erased received the same treatment, or whether certain cohorts of data were quietly preserved for training or evaluation purposes while others were actually purged.
For anyone who has used ChatGPT and later requested deletion of their conversation history, the allegation carries direct personal stakes. OpenAI’s privacy policy and public statements have long emphasized user control over stored conversations. A court finding that tens of millions of logs persisted in a company database after deletion claims would erode the credibility of those assurances and could invite regulatory scrutiny beyond the courtroom.
Court filings and deposition evidence behind the 78 million figure
The 78 million figure comes from the publishers’ sanctions motion itself, which cites deposition testimony and internal exhibits produced during the litigation. The motion, docketed in case records for the Southern District, describes a de-identified reservoir that OpenAI maintained even as it told the plaintiffs the data was no longer available. The publishers contend that this reservoir contained conversation logs that could shed light on how ChatGPT ingested and reproduced copyrighted material.
Bloomberg Law reported that The New York Times is seeking sanctions against OpenAI in the same copyright case, confirming that the motion targets alleged discovery misconduct rather than a standalone privacy complaint. The publishers’ legal theory ties the retained logs directly to the question of whether OpenAI trained its models on copyrighted text: if the company kept millions of user conversations that included publisher content, those logs could serve as evidence of how training data flowed through ChatGPT’s systems.
The distinction between “de-identified” and “deleted” is central to the dispute. De-identification strips personal identifiers such as names and email addresses but preserves the substance of the conversations. From a copyright plaintiff’s perspective, the content of the chats, not the identity of the users, is the relevant evidence. OpenAI’s decision to retain the logs in de-identified form rather than destroy them outright suggests the company saw ongoing value in the data, even as it represented to opposing counsel that the records were gone.
Appellate and related federal court dockets trace the procedural history of the case, which began when The New York Times sued OpenAI and Microsoft in late 2023. The sanctions motion is the latest escalation in litigation that has already produced extensive discovery battles over what training data OpenAI used and how it stored or discarded user interactions. The emerging record shows a widening conflict not only over the underlying copyright issues but also over how far a technology company must go to preserve and disclose internal data once it is sued.
Open questions about OpenAI’s deletion practices and what to watch next
Several gaps in the public record prevent a full accounting of what happened. The deposition transcripts and exhibits cited in the sanctions motion remain sealed or heavily redacted, so the exact timeline of OpenAI’s deletion commitments is not visible outside the courtroom. No primary data from OpenAI or any regulator shows how many of the 78 million logs were actually used in model training versus stored for other purposes such as safety evaluation or system monitoring. That uncertainty makes it difficult to gauge whether the alleged retention primarily affects copyright claims, user privacy concerns, or both.
OpenAI has not released a public statement addressing the specific allegation in the sanctions motion. Without that response, it is unclear whether the company will argue that de-identification satisfied its obligations, that the logs were preserved under a litigation hold, or that the plaintiffs mischaracterize the facts. Each defense would carry different implications for users and for the broader copyright case. If OpenAI leans on de-identification, the dispute will likely focus on what its prior statements to the court and to users actually promised. If it invokes a litigation hold, the question will be whether the company clearly explained that obligation when it discussed deletion practices with the plaintiffs.
The question of differential treatment across user groups also remains unanswered. If OpenAI retained 78 million conversation logs while processing deletion requests from individual users during the same period, any measurable gap between what was promised and what was kept could become a separate regulatory issue. Privacy authorities in the European Union and several U.S. states have active investigations or rulemaking processes around data minimization and user control, and a judicial finding that large reservoirs of “deleted” chats persisted in practice would likely feed into those efforts. Even without a formal investigation, consumer advocates are likely to seize on the case as a test of whether AI companies are living up to their public commitments.
For now, the immediate stakes are procedural but significant. Sanctions could reshape the evidentiary landscape of the Times lawsuit, strengthening the publishers’ leverage in settlement talks or at trial. An adverse inference instruction, for example, would allow a jury to assume the missing or mishandled data would have been unfavorable to OpenAI, a powerful tool in a case that turns on opaque training pipelines and proprietary datasets. Monetary penalties would be more symbolic but could still signal judicial impatience with how the company has handled discovery.
Beyond this case, the dispute over the 78 million logs highlights a broader tension in AI development: companies want to retain large volumes of interaction data to refine their models, but courts and regulators are increasingly skeptical of expansive data retention justified only by innovation. How judges in the Southern District of New York respond to the publishers’ motion will be closely watched by other AI firms facing similar lawsuits. A ruling that treats de-identified logs as effectively “retained” for legal purposes, despite public assurances of deletion, would set a precedent that could force companies to rethink their data lifecycles from the ground up.
Users, meanwhile, have limited visibility into these internal choices. The sanctions motion offers a rare glimpse into how at least one major AI provider handled the tension between promising deletion and preserving data for technical and legal reasons. Whatever the court ultimately decides, the episode underscores the importance of precise language in privacy policies, transparent explanations of retention practices, and robust oversight of how those promises are implemented once litigation and regulatory scrutiny begin.
More from Morning Overview
*This article was researched with the help of AI, with human editors creating the final content.