AIO APEX

Anthropic and OpenAI agree to embed independent safety evaluators inside their labs

TechCrunch
Share:
Anthropic and OpenAI agree to embed independent safety evaluators inside their labs

Anthropic CEO Dario Amodei has proposed a new model for AI oversight: independent third-party evaluators permanently embedded inside frontier AI laboratories, with the same access as employees, company badges, and the right to publish their findings. OpenAI CEO Sam Altman has pledged that OpenAI will match the commitment.

What the Proposal Entails

Under Amodei's framework, each frontier AI company would give an ongoing team of outside evaluators — organizations like METR — the kind of access typically reserved for senior employees: the ability to examine training pipelines, safety practices, and model alignment processes before deployment, not just after. The arrangement is explicitly designed to replace periodic, point-in-time audits with continuous monitoring.

Amodei cited banking regulation as precedent. Some financial regulators embed supervisors directly within firms, sitting alongside employees throughout the year rather than conducting occasional external reviews. The proposal adapts that model to AI development, where the risks of missing a safety issue during training can be irreversible.

Three-Part Governance Framework

The evaluators proposal is one part of a broader framework Amodei outlined in a September 12 essay. The other two components are: common safety standards among democratic nations' frontier AI companies, creating a coordinated floor below which no company can fall; and international coordination — including formal agreements with China and other governments — to prevent races to the bottom on safety globally.

Altman confirmed OpenAI's agreement to the evaluators commitment, saying: "I agree with Dario that we need to pace the frontier." OpenAI specified that "pacing" means deliberately slowing some aspects of development to allow safety measures to keep up — not stopping development entirely.

The Independence Problem

Researchers and AI safety advocates welcome the unprecedented access Amodei has described. But the proposal, as currently structured, gives evaluators extraordinary visibility with limited formal authority. Under the proposed arrangement, evaluators could examine systems, report incidents, and publish findings — but they would have no veto power over training runs or model releases. Acting on their findings would remain at the companies' discretion.

CNBC called the evaluators "neutral watchdogs" while noting that without statutory enforcement mechanisms, the arrangement depends entirely on the companies choosing to act on what evaluators surface. Critics describe the risk as legitimacy-washing: using the appearance of independent oversight to satisfy critics while retaining full control over what happens next.

Congressional Movement in Parallel

A bipartisan House proposal is moving through Congress that would mandate embedded evaluators at frontier AI companies rather than leaving it to voluntary commitment. The bill would give evaluators formal legal standing and reporting requirements — addressing the independence gap the voluntary proposal leaves open.

Why This Matters

Amodei has argued publicly this month that AI could outrun our ability to understand and control it within six to twelve months. If that timeline is credible, the question of how independent oversight actually works — not in principle but in practice — becomes urgent. The embedded evaluator model represents the most access any external organization has received to the internal workings of a frontier AI lab. Whether it translates into genuine accountability or becomes the kind of audit theater the banking analogy was meant to avoid will be determined by what evaluators are actually allowed to see, say, and act on, as first reported by TechCrunch.

Originally reported by TechCrunch. Read the original article for additional details.

View original source
Share: