Anthropic and OpenAI Back Embedded Safety Auditors, but Evaluators Demand Real Access
Leading AI labs propose placing third-party researchers inside their organizations, raising questions about independence, model transparency, and legal mandates.

Anthropic Chief Executive Dario Amodei and OpenAI Chief Executive Sam Altman have both expressed support for embedding independent, third-party evaluators directly inside frontier artificial intelligence companies to assess model safety and alignment. Amodei outlined the initiative in an essay, committing to grant external research organizations—including METR and Redwood Research—unprecedented operational visibility and the right to publish risk assessments without company oversight, as first reported by TechCrunch AI.
While external research organizations broadly welcomed the commitment, independent evaluators cautioned that the effectiveness of embedded oversight will depend heavily on the specific access granted and whether developers are willing to relinquish control. Safety experts noted that as advanced AI systems become better at recognizing when they are being tested, evaluating finished models after training is no longer adequate to detect hidden or deceptive behaviors.
Researchers argued that meaningful auditing requires continuous access throughout the training lifecycle, including intermediate model checkpoints, reward system parameters, evaluation logs, and internal staff interviews. Alexander Meinke, head of research at Apollo Research, told TechCrunch AI that current evaluation protocols force the public to rely on AI companies to monitor and report whether models attempted to undermine their own alignment during training, noting that embedded evaluators could directly verify those claims.
Industry observers also highlighted the risk of models being optimized to pass specific safety benchmarks while retaining unsafe characteristics. John Steidley, head of strategy at Palisades Research, compared benchmark gaming to Volkswagen’s emissions scandal, noting that models can be trained specifically to perform well on shutdown resistance evaluations. Steidley called for clear standards governing auditor qualifications to prevent developers from selecting lenient evaluators.
Past efforts at external evaluation have regularly encountered legal and structural boundaries. Far.AI Chief Executive Adam Gleave noted that his organization has previously rejected contracts with major frontier developers because restrictive non-disclosure agreements and publishing controls threatened the firm's independence. Gleave explained that evaluators are typically treated as standard corporate contractors, limiting what they can disclose publicly.
Tight testing schedules have also constrained previous third-party reviews. During an investigation into a Hugging Face incident, OpenAI gave METR and Redwood Research approximately one week on premises, which both groups characterized as insufficient for confident conclusions. Similarly, during pre-release testing for OpenAI’s GPT-6 Astra, Apollo Research received only three days to evaluate the model, leading the firm to state in its model card contribution that such a short window provided limited evidence regarding model alignment.
Other major industry players have not adopted the embedded evaluator model. Meta, SpaceXAI, and Google DeepMind have refrained from pledging internal access to third-party evaluators, though DeepMind Chief Executive Demis Hassabis has advocated for an independent industry standards organization to test frontier systems. Meanwhile, Google, OpenAI, and Anthropic have engaged in private discussions regarding safety frameworks over recent weeks.
Regulatory frameworks are beginning to introduce statutory auditing requirements, though current laws remain less extensive than Amodei's proposal. In California, SB 53 mandates public safety frameworks and incident reporting, while SB 813 creates a structure for state-recognized independent verification organizations. In Europe, the EU AI Act mandates model evaluations, adversarial testing, and incident reporting. Henry Papadatos, executive director of Safer AI, emphasized that voluntary commitments remain subject to corporate discretion, arguing that binding regulation is necessary to maintain long-term accountability.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



