Anthropic and OpenAI Push for Embedded AI Watchdogs, Drawing Scrutiny From Policy Experts
Proposals to place third-party risk evaluators inside frontier AI labs face criticism over their lack of enforcement power and potential conflicts of interest.

Anthropic and OpenAI have recently advocated for placing independent third-party risk evaluators inside their facilities to monitor high-stakes artificial intelligence development. The proposal, initially detailed in an essay by Anthropic Chief Executive Dario Amodei, aims to mitigate safety risks associated with rapidly advancing frontier models. However, legal scholars, industry experts, and auditors are raising concerns that the initiative lacks regulatory teeth and could amount to corporate self-policing, as reported by CNBC Business.
Amodei's recommendation followed the departure of former Anthropic researcher Jacob Coxon, who stepped down after cautioning that developers are building systems they may ultimately fail to control. In his essay, which has sparked debates across Washington and the technology sector, Amodei likened embedded evaluators to regulatory supervisors used in the banking sector. Under his plan, Anthropic would grant external evaluators access equivalent to internal risk departments and permit them to publish findings independently, subject only to minor redactions.
Legal experts quickly pointed out key distinctions between banking oversight and the proposed AI model. Julie Andersen Hill, dean of the University of Wyoming College of Law, told CNBC Business that government bank examiners possess sweeping legal enforcement powers—including the authority to suspend operations, replace executives, or liquidate failing institutions. By contrast, neither Anthropic nor OpenAI has proposed granting external reviewers formal power to halt training runs or block product releases, rendering the comparison fundamentally flawed in Hill's view.
Firms currently involved in pre-release model assessments report a gap between existential fears and current technical reality. Albert Ziegler, head of AI at cybersecurity vendor XBOW, stated that his firm routinely receives early access to unreleased software from major developers including OpenAI and Anthropic to conduct evaluations in isolated environments. While Amodei warned that autonomous agent swarms could potentially disrupt internet infrastructure and inflict hundreds of billions of dollars in damages within six to twelve months, Ziegler noted that XBOW has primarily uncovered minor formatting errors and routine safety triggers rather than insidious systemic threats.
Questions have also emerged surrounding evaluator independence and industry ties. Amodei identified the non-profit organization Model Evaluation and Threat Research (METR) as a potential embedded partner. METR recently brought on former Anthropic researcher Joe Benton to lead embedded risk assessments and was tasked with reviewing cybersecurity evaluation incidents involving Anthropic's Claude model. While METR states that it does not accept direct funding from AI companies, the organization acknowledged in its own risk reports that staff members share social connections and facility space with lab researchers.
Critical voices emphasize that meaningful oversight requires external governance rather than voluntary arrangements. Deborah Raji, an AI accountability researcher at the University of California, Berkeley, noted that true audit independence mandates strict rules governing conflicts of interest and third-party qualification, preventing companies from selecting their own assessors. Addressing safety concerns at the Politico Decoded summit in Washington, Anthropic Head of Public Policy Sarah Heck acknowledged that tech firms cannot rely on an "honor code," confirming that Anthropic communicates daily with the White House and maintains ongoing discussions with Congress.
Beyond conflict-of-interest concerns, experts point to a shortage of qualified evaluators and a lack of standardized legal rules for frontier AI. Christina Ho, chief assurance officer at accounting firm Oath and former board member of the Public Company Accounting Oversight Board, observed that traditional financial audits benefit from established operational standards, whereas AI verification requires rare technical expertise to evaluate live systems and outputs. Without statutory authority or standardized guidelines, legal scholars like Hill argue that voluntary evaluator access cannot deliver genuine accountability without an independent mechanism capable of stopping unsafe deployments.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



