Independent AI safety evaluators have moved from the industry's margins to the center of a debate over who guards the guarders. Companies like Anthropic and OpenAI now face pressure to embed third-party assessors as federal regulation remains absent, but the emerging ecosystem is underfunded, understaffed and structurally compromised by its reliance on the labs it evaluates.
The evaluator sector consists mostly of small nonprofits and startups including Model Evaluation and Threat Research (METR), Apollo Research, Transluce and Vals AI. Their role involves assessing AI model capabilities, flagging safety risks and documenting when systems misbehave. Anthropic CEO Dario Amodei pledged last month to embed independent evaluators into his company. OpenAI CEO Sam Altman endorsed the approach. President Trump and major U.S. tech companies backed the idea in September through a voluntary accord.
But fundamental questions remain unanswered: who funds these groups, what access will they receive and what reporting structures will govern their work?
METR raised $71 million in commitments over six months, up from $13.6 million in total 2024 contributions, according to its IRS filing. Vals AI announced a $40 million funding round in August and grew from eight employees to roughly 30 this year. Yet Kevin Werbach, faculty director of the Wharton Accountable AI Lab, calls the ecosystem "not robust enough right now." METR employs fewer than 50 full-time staff.
Suresh Venkatasubramanian, a computer science professor at Brown University, told CNBC: "To a degree, the problem, as always, is money. Who is paying for these companies to do their work? How are they going to support them? You need an ecosystem, you need a viable business model for this."
Structural conflicts have already emerged. OpenAI fired three employees last week for "violating our policies on accessing and handling sensitive company information," according to a spokesperson. Two of those employees, Mikita Balesni and Tomek Korbak, said they believed they were dismissed because of how they communicated with third-party evaluators. Balesni wrote on X that colleagues feared retaliation for speaking with external parties. OpenAI disputed that claim and said it is "actively finalizing contracts with third-party safety assessors" and remains "committed to embedding external assessors."
The power imbalance is stark. Labs have raised tens of billions of dollars and employ thousands. Evaluators remain financially dependent on those same companies. Venkatasubramanian noted: "If you want true third-party evaluation, you need true independence financially and otherwise."
