Skip to content
Breaking:

OpenAI Outlines Priorities and Ground Rules for External AI Safety Audits

A framework document details four assessment domains and seven operational principles for independent safety testers.

By The Company Wire3 min read
Share
OpenAI — OpenAI Outlines Priorities and Ground Rules for External AI Safety Audits
OpenAI — OpenAI Outlines Priorities and Ground Rules for External AI Safety Audits. Photo: The Next Web.

OpenAI published a framework on Tuesday establishing four primary assessment priorities and seven operating principles for independent safety evaluators examining its artificial intelligence models. Authored by Lama Ahmad, who leads OpenAI's engagements with external safety experts, the document provides operational detail following the company's announcement earlier the same day that it would open models to outside scrutiny earlier in their development cycles.

The framework centers on safety cases, structured arguments supported by testable claims demonstrating how the lab manages risk across training, evaluation, internal deployment, and external deployment. The document requires safety cases to disclose underlying assumptions, uncertainties, and residual risks. The approach reflects recent statements from OpenAI Chief Executive Sam Altman, who proposed a federal oversight system grounded in safety cases to pace rather than halt model development.

The first assessment area focuses on independently evaluating these safety cases, testing whether evidence holds up, whether teams followed specified conditions, and whether training regimes inadvertently reward deception, hacking, or restriction circumvention. The second covers the safeguard stack, providing assessors with "grey box access" to evaluate resilience against jailbreaks and capability uplift in biological and cyber domains. Testers are also tasked with probing whether misalignment monitors contain gaps that could cause a loss of control, and how reliably chain-of-thought monitoring performs as models advance.

The third priority addresses capability evaluations under the company's Preparedness Framework, which tracks risks across cybersecurity, chemical and biological domains, and AI self-improvement. The document asks evaluators to examine whether safety thresholds are correctly calibrated and refreshed as models exceed them. OpenAI disbanded its dedicated Preparedness team in August, reassigning its responsibilities across other internal groups. The fourth domain involves independent post-incident investigations of misalignment events, citing an earlier incident on Hugging Face as an example requiring cyber forensics and chain-of-thought analysis at scale.

To govern third-party reviews, OpenAI outlined seven operational principles. Assessors and the lab must agree on scope and pre-register safety claims before evaluations begin, noting whether claims originated from the company or the auditor. Testers must disclose their methodologies, acknowledge uncertainties, separate direct findings from interpretations, and disclose conflicts of interest, including financial ties or past involvement. Assessors are to receive access proportionate to their claims within legal, security, and intellectual property constraints, using company representatives or privacy-preserving tools when direct access is impractical.

Three provisions offer specific protections for the company. If assessors cannot satisfy security standards in their own environments, testing may be restricted to OpenAI-managed devices or premises. The guidelines also state that developers should receive a reasonable, though unspecified, period to remediate identified issues before publication. In addition, labs may request redactions of sensitive data, though evaluators retain editorial independence and may disclose where redactions occurred and how they altered the final report.

As reported by The Next Web , the document does not address who pays for the audits, leaving financial models unresolved even as legislative measures such as California's Senate Bill 813 establish frameworks for independent verification organizations. OpenAI said it is in active discussions with multiple unnamed third parties regarding assessment proposals, noting that audits will be launch-agnostic and may span from several weeks to several months.

Sources

  1. The Next Web

Company: OpenAI

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.