Skip to content
Breaking:

Unsealed Court Filings Reveal Microsoft and OpenAI Internal Warnings Over Web Data Scraping

Court documents from The New York Times' copyright lawsuit show internal concerns at Microsoft and OpenAI over the impact of AI models on digital publishers.

By The Company Wire4 min read
Share
OpenAI — Unsealed Court Filings Reveal Microsoft and OpenAI Internal Warnings Over Web Data Scraping
OpenAI — Unsealed Court Filings Reveal Microsoft and OpenAI Internal Warnings Over Web Data Scraping. Photo: The Verge.

Newly unsealed court documents from The New York Times' copyright lawsuit against OpenAI and Microsoft reveal that internal teams at both tech companies raised early concerns about the consequences of using scraped web data to train generative artificial intelligence models. As reported by The Verge (https://www.theverge.com/ai-artificial-intelligence/997633/openai-microsoft-chatgpt-ai-new-york-times-doom-loop-theft-google-zero), the 92-page filing includes internal communications, testimony, and strategy papers acknowledging that AI assistants could disrupt the broader economic model of online publishing by replacing direct web traffic.

Internal commentary from Microsoft personnel highlights friction within the organization over data acquisition practices. Brent Hecht, director of applied science at Microsoft, privately described the widespread harvesting of online data for models like ChatGPT and Copilot as the largest theft of labor in human history and argued that corporate legal defenses made a mockery of fair use. Responding to the filing, Microsoft spokesperson Alex Haurek told The Verge that Hecht's remarks reflected an individual viewpoint rather than an official corporate position or legal analysis. In addition, Jordan Usdan, general manager for data strategy and operations at Microsoft AI, stated in court records that Hecht was hired to offer speculative, academic perspectives and does not speak on behalf of the company regarding content creator impacts.

Internal strategy documents from Microsoft cited in the filing warned that the company's content approach had triggered a "doom loop" capable of harming both model accuracy and the wider web ecosystem. The records noted that it was highly unusual for an end product to undermine the financial health of its primary data suppliers, adding that large language models act as products that fundamentally erode their own content supply chains. Microsoft documents further observed that creators never intended for their content to be utilized in this manner without financial compensation.

The filings also detailed executive perspectives from Microsoft Chief Executive Officer Satya Nadella and OpenAI co-founder Greg Brockman. Nadella acknowledged that conversational tools have effectively replaced search queries by removing the requirement for users to navigate to source websites. While Nadella was later quoted stating that paywalled material should be licensed, an OpenAI representative admitted in court papers that he was unaware of internal mechanisms designed to identify or strip paywalled data from training sets. Addressing Nadella's statements, Microsoft spokesperson Haurek told The Verge that the chief executive was discussing broad shifts in how people consume information rather than making legal concessions on copyright questions.

OpenAI's internal assessments similarly noted risks regarding text reproduction and data memorization. Despite acknowledging that preventing memorization was necessary to minimize copyright violations, internal discussions indicated that GPT-4 had absorbed large amounts of training data, making the system prone to exact regurgitation. The court filing cited instances where ChatGPT generated long, verbatim passages from published articles originating from outlets such as The New York Times, The Denver Post, the Mercury News, Lifehacker, and Eurogamer.

Additional quotes in the document illustrate how OpenAI staff viewed the chatbot's potential to displace online media outlets. OpenAI Policy Director Jack Clark noted internally that the company was constructing tools that substitute for human labor in cultural production, while separate internal files referred to ChatGPT as a modern newsstand. Nick Turley, head of ChatGPT at OpenAI, observed that users have no good reason to click through to primary web sources once a direct answer is rendered by the model. Furthermore, economic and media specialists at OpenAI speculated that search referrals to news outlets could drop by as much as 60 percent due to AI-driven search summaries such as Google's AI Overviews.

The unsealed court records add further context to the high-stakes legal struggle over whether scraping copyrighted material to train commercial artificial intelligence systems constitutes fair use under federal law. Media companies continue to contend that tech firms are building commercial products using proprietary content without authorization, while OpenAI and Microsoft maintain that their model development practices conform to established legal principles.

Sources

  1. The Verge

Company: OpenAI

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.