AI Sources
Primary sourceAIArticle··5 min read

Anthropic Opens Its Doors to Internal Evaluators Backed by $1 Billion

Accenture will place security auditors inside Anthropic, with access comparable to that of an employee — and billed to the lab itself.

Anthropic Opens Its Doors to Internal Evaluators Backed by $1 Billion
Source : Anthropic · anthropic.comView original

In brief

Anthropic announces a partnership with Accenture for "embedded evaluation": independent evaluators working from inside the lab, with employee-level access, to audit models, run red-teaming exercises and verify safety commitments. Each party plans to invest at least $1 billion over five years. The arrangement is unprecedented, its rules aren't written down yet, and for now Anthropic is directly funding the work of its own auditor.

🍺 Bar-stool version

Anthropic has decided to install inspectors inside its own walls, badge and all, with access to models still in training and the right to talk to any employee they want. The funny detail is that the inspector in question, Accenture, is paid by Anthropic — a bit like choosing and paying for your own car's inspection technician. The lab says it plainly: there's no standard today, no public funding, so they're working with what they've got in the meantime. It's shaky, but it's the first time a frontier lab has let an outsider watch how the sausage gets made, and that beats the alternative.

Key takeaways

  1. 1

    Anthropic partners with Accenture for "embedded" evaluation of its frontier models, a commitment stemming from the CEO's essay "We Must Pace the Frontier."

  2. 2

    The work will be led by Faculty, Accenture's specialized AI unit: model evaluation and red-teaming, alignment assessments, and guardrail testing.

  3. 3

    Anthropic and Accenture each plan to invest at least $1 billion over five years to build out this capability.

  4. 4

    Unlike current external evaluators, embedded evaluators will have employee-comparable access: tracking training, deployment decisions, and direct discussions with teams.

  5. 5

    No standard yet exists for what these evaluators should have access to or how they should publish their findings.

  6. 6

    Anthropic is directly funding Accenture's work, for lack of pooled or public funding; the lab is also in talks with METR and other nonprofit evaluators, which are funded independently.

  7. 7

    The partnership is non-exclusive: other evaluators will be announced in the coming weeks, and Accenture will work with other AI developers too.

What changes with "embedded" evaluation

Until now, frontier model audits have been done from the outside: an organization gets controlled access to a model, often late and often partial, and produces a report. Anthropic is proposing something different: placing evaluators inside, with access "comparable to that of an employee."

Concretely, this means watching models take shape during training, following the decisions that govern their construction and deployment, and speaking directly with the people making them. The object of the audit is no longer just the finished model — it's the process leading to it.

The stated scope covers model evaluation and red-teaming, alignment assessments, and guardrail testing. All of it will be led by Faculty, Accenture's specialized AI unit.

Anthropic stresses one point: this setup doesn't dilute its responsibility. Model safety remains its own; the embedded evaluator serves to make that responsibility verifiable rather than transferring it.

Why Accenture, and why a billion dollars

The choice may come as a surprise: Accenture isn't an AI safety research lab, it's a consulting giant deploying AI for businesses and governments. Anthropic frames this as an asset: this real-world enterprise deployment knowledge feeds a safety approach grounded in practice rather than benchmarks.

The announced figure is massive: each party plans to invest at least $1 billion over five years to build this capability. This is the main signal of the announcement — independent evaluation stops being a symbolic budget line and becomes infrastructure to be funded.

Note that this sum funds the building of an evaluation capability, not a standard service contract. It requires recruiting, training and equipping teams capable of tracking frontier models in real time — a job that barely exists today.

The funding knot

Anthropic doesn't hide the problem: the lab is paying for its own auditor. The justification is pragmatic — there's currently no pooled fund or public financing for this kind of work, and urgency doesn't allow waiting.

In the long run, Anthropic says it wants funding from pooled or public sources, a position it already laid out in its Advanced AI Framework published in June. In the meantime, the lab says it's working with different evaluators under different financial arrangements.

That's the point of the discussions mentioned with METR and other nonprofits: piloting elements of embedded evaluation with organizations funded from their own resources, and thus structurally less dependent on Anthropic.

One blind spot remains: nothing is said about what an embedded evaluator is allowed to publish, or what happens in case of a fundamental disagreement. Anthropic explicitly acknowledges this — standards for access and reporting don't exist yet.

Toward an ecosystem of evaluators

Anthropic frames the partnership as non-exclusive in both directions: other evaluators will work with the lab, to be announced in the coming weeks, and Accenture will take on similar roles with other AI developers.

The stated ambition is an ecosystem: multiple organizations auditing multiple labs, under shared standards. A single evaluator per lab would be both a point of failure and a risk of capture.

Context adds weight to the announcement. Anthropic recently reported three incidents in which Claude models gained unauthorized access to real computer systems, and was already planning an independent review with METR. Embedded evaluation is the structural response to this kind of episode.

Independent embedded evaluators do not reduce our accountability, but help to make it more verifiable.
There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find.
Ultimately, we believe frontier AI needs an ecosystem of evaluators operating with shared standards.

Why it matters

This is the first serious attempt by a frontier lab to make its internal practices verifiable by a third party, rather than merely documented in a report released alongside the model. The move is real: a billion dollars over five years, employee-level access, an explicit admission that standards are missing. But the structure remains fragile on two points. First, funding: an auditor paid by the audited party, however good the intentions, doesn't have the independence of a regulator — Anthropic says so itself and calls for public or pooled funds, which amounts to admitting the current solution is a stopgap. Second, the choice of Accenture, a firm that otherwise sells AI deployment to companies and governments, and is therefore a stakeholder in the very market it's meant to oversee. The real test will come the day an embedded evaluator wants to publish something inconvenient, and we find out whether it's allowed to.

#ai#anthropic#safety#governance#evaluation
Original source
Partnering with Accenture on embedded evaluation
Anthropic
Open the article

Read next