Skip to content
Breaking:

Unredacted Filings Show Microsoft and OpenAI Executives Discussed AI Scraping as Content Theft

Newly unsealed court documents in The New York Times lawsuit reveal internal concerns over publisher market substitution, paywall circumvention, and traffic drops.

By The Company Wire4 min read
Share
Microsoft — Unredacted Filings Show Microsoft and OpenAI Executives Discussed AI Scraping as Content Theft
Microsoft — Unredacted Filings Show Microsoft and OpenAI Executives Discussed AI Scraping as Content Theft. Photo: TechCrunch AI.

Newly unredacted legal filings in the ongoing copyright infringement lawsuit brought by The New York Times against Microsoft and OpenAI reveal that executives at both companies privately acknowledged that AI data scraping practices severely threatened news publishers, with one senior technical leader characterizing the process as systemic theft.

The unsealed court brief, details of which were first reported by TechCrunch AI, incorporates internal records and deposition testimony gathered during discovery. In a January 2023 internal document, Microsoft Director of Applied Science Brent Hecht characterized AI training data collection as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." A subsequent January 2024 internal presentation authored by Hecht documented that Microsoft's Copilot search feature caused click-through rates for The New York Times website to fall by up to 93 percent compared to traditional Bing search, creating what he described as a "doom loop" that threatened both the open web and the future performance of AI models.

In the same 2024 presentation, Hecht observed that it was "highly unusual that an end-product threatens the economic foundations of its essential suppliers," noting that Microsoft had created that precise dynamic within its commercial language model content supply chain. The New York Times brief argues that these internal acknowledgments contradict the fair use defense raised by both technology companies, which requires showing that a transformative work does not displace or harm the primary market for the original material.

Statements from OpenAI leadership cited in the filing also highlighted market substitution risks. Nick Turley, the head of ChatGPT at OpenAI, wrote in internal communications that digital publishers faced an "existential threat" from conversational AI platforms that act as "largely substitutive" products that would become increasingly so as underlying technology improved. OpenAI President Greg Brockman described the models as "excellent at news," while Microsoft Chief Executive Officer Satya Nadella acknowledged during a deposition earlier this year that conversing with AI platforms replaces the need for users to navigate to source websites.

During his deposition, Nadella testified that any material situated behind a paywall should require formal licensing by any entity seeking to use it for model grounding or training. Nadella further stated under oath that if he had known OpenAI was scraping paywalled publications, he would have exercised Microsoft's contractual authority to demand that OpenAI retrain its systems without that data.

The unredacted filing also discloses the scale of content used to build the generative AI models and details joint dataset sharing between the two corporate partners. OpenAI's mid-training datasets included more than 91,692 copies of works published by The New York Times, the Daily News, and the Center for Investigative Reporting, while a dataset sourced from Common Crawl contained over 2 million documents from nytimes.com alone. Additionally, OpenAI shared its full GPT-3 training dataset with Microsoft, while Microsoft transferred content to OpenAI through internal programs known as Project Taxi and Project Mango, the latter yielding a dataset containing at least 160,903 distinct published news works.

According to the plaintiffs' brief, OpenAI personnel actively sought techniques to bypass publisher subscription barriers. When OpenAI researcher Nick Ryder messaged Brockman regarding a method to circumvent The New York Times paywall, Brockman replied, "ah nice." The filing also alleges that OpenAI compiled custom datasets, including WebText and WebText2, that leaned heavily on news content, and deliberately removed copyright notices from training data to prevent models from outputting copyright disclaimers to end users.

The unredacted brief represents a major development in the three-year-old litigation. While federal courts evaluating generative AI lawsuits have largely favored technology providers under fair use standards—supported recently by a brief filed by the Trump administration backing OpenAI's unlicensed training practices—the newly unsealed internal statements present detailed evidence regarding how key industry figures assessed the legal and economic impact of their data gathering strategies.

Sources

  1. TechCrunch AI

Company: Microsoft

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.