In July 2026, The New York Times and The Daily News escalated their copyright infringement lawsuit against OpenAI by filing a motion for sanctions. The publishers allege that OpenAI concealed evidence contradicting the company's previous claims about technical limitations in searching its own data.

According to an April deposition by OpenAI data privacy engineer Vinnie Monaco, the company had already conducted internal searches of its training corpus to identify copyrighted journalism. More significantly, OpenAI maintained a database of approximately 78 million de-identified ChatGPT conversations used internally to assess the extent of copyright infringement in its outputs. Additionally, the company had implemented a "Bloom" filter as part of "Project Giraffe," a suite of tools designed to detect and log instances where ChatGPT reproduced existing content, deployed shortly after the lawsuit commenced.

Throughout the two-year litigation, OpenAI had insisted it lacked the technical capability to search its training data and argued that retrieving, processing, and de-identifying its massive conversation logs would be operationally burdensome and raise privacy concerns. The publishers sought this data to verify whether their copyrighted journalism appeared in OpenAI's training dataset and to determine how frequently ChatGPT generated responses containing or reproducing their content.

The discovery of these tools and datasets represents a significant contradiction of OpenAI's prior legal arguments. When the publishers originally requested 120 million chat logs, OpenAI negotiated the sample down to 20 million. The company eventually submitted that reduced sample to the court in December 2025, but it contained extensive redactions that rendered it effectively unusable according to the court's assessment.