Skip to content
Breaking:

OpenAI Says It Cannot Rule Out Use of De-Identified Data from Specific Users in Model Training

The AI company described the outcome as unlikely but acknowledged the technical impossibility of definitively excluding anonymized usage data.

By The Company Wire3 min read
Share
OpenAI — OpenAI Says It Cannot Rule Out Use of De-Identified Data from Specific Users in Model Training
OpenAI — OpenAI Says It Cannot Rule Out Use of De-Identified Data from Specific Users in Model Training. Photo: Techmeme.

OpenAI has acknowledged that it is unable to completely exclude the possibility that de-identified data originating from product interactions by users identified as Buckmaster and Alpöge was utilized to train and improve its artificial intelligence models.

The artificial intelligence organization characterized such an outcome as improbable, according to disclosures first reported by Techmeme on September 8, 2026. Nevertheless, OpenAI confirmed that it "cannot rule out that de-identified data derived" from the two individuals' use of its services ultimately fed into its model refinement processes.

The statement points to the technical hurdles inherent in tracking data provenance across large language model training pipelines. When user prompts and interactions are stripped of personally identifiable information to create anonymized datasets, establishing the exact origin of specific data points becomes difficult to verify or disprove after the fact.

In addressing the activity of Buckmaster and Alpöge, OpenAI noted that its evaluation of product usage records leaves open the potential that anonymized inputs were absorbed into training corpora. Standard privacy and data-handling workflows frequently anonymize user text to train future iterations of software, making total isolation of specific accounts challenging once de-identification has occurred.

The disclosure highlights ongoing industry questions surrounding data governance, user privacy, and transparency in machine learning. As technology firms process vast quantities of user-generated information to optimize generative systems, the distinction between active account data and anonymized training sets remains a sensitive focus for developers and users alike.

OpenAI maintained that the actual incorporation of data from Buckmaster and Alpöge remains unlikely. However, by publicly conceding that the possibility cannot be ruled out, the company underscored the reality that de-identified data streams can persist within the broad architecture of modern AI development.

Sources

  1. Techmeme

Company: OpenAI

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.