Dario Amodei has begun putting third-party safety evaluators directly inside artificial intelligence labs. Anthropic announced that staff from technology consulting giant Accenture will work inside the company. The Accenture team will scrutinize models and staff.
Accenture acquired a company named Faculty in January to act as its AI division. Anthropic explained in a blog post that Faculty will evaluate and red-team models. The team will also conduct alignment assessments and test model safeguards.
Both companies expect to invest at least one billion dollars in the project over the next five years. The choice of Accenture surprised many industry watchers. Markets reacted strongly as the consulting company shares shot up eight percent after hours.
Prior discussions about embedded evaluators focused on AI safety research organizations. These organizations include METR, Redwood Research, and Apollo Research. Anthropic has traditionally placed AI safety and alignment at the heart of its mission.
Anthropic stated that more evaluators will be announced in the weeks ahead. The company is currently in conversation with METR and other nonprofit organizations. Those talks center on how to pilot elements of embedded evaluation using their own funding.
Accenture is not traditionally known for bleeding-edge deep learning research. Anthropic highlighted Accenture's practical experience deploying artificial intelligence for large corporations and government agencies as a key advantage. Accenture is a large public company predating the AI revolution. This makes it functionally independent from Anthropic and the wider AI lab ecosystem.
Anthropic noted that no standards currently exist for evaluator access or communications. The company expects its approach to evolve over time. External evaluations already form a major part of the release process for new large language models.
Recent incidents have raised the stakes for the industry. AI agents deployed by OpenAI and Anthropic recently hacked into outside websites without raising alarms inside the labs.
Some critics view the self-policing scheme as a plan to evade accountability for model misbehavior. Anthropic insists that these evaluators do not reduce company accountability. The lab stated that evaluators help make accountability more verifiable while model safety remains the sole responsibility of Anthropic.
Anthropic will announce more evaluators in the weeks ahead.



