New Issue: Orbital Catastrophe Ahead? Read Now

AI leaders want to slow frontier AI. Who will make sure they really do?

Independent auditors are emerging as a key part of plans to rein in frontier AI. But the systems are evolving faster than the methods used to evaluate them

A woman with a clipboard stands beside a mainframe computer under a sign reading SYSTEM ON-LINE.

Exactly how AI auditors should scrutinize the world’s most powerful models remains unsettled.

Lambert/Hulton Archive/Getty Images

On September 12 Anthropic CEO Dario Amodei published an essay urging the industry to “slow the pace at which we improve the capabilities of AI models.” To make such a slowdown verifiable, he proposed embedding outside evaluators inside frontier AI companies. At Anthropic, he wrote, such AI auditors would get access comparable to the company’s own risk teams and the right to publish their findings.

His idea puts enormous weight on a field whose rules are still being written. More and more, AI auditors are being asked to serve as independent checks on the companies that are building the most powerful models, even though no standard exists for what a rigorous evaluation entails. Amodei’s proposal would add their most high-profile responsibility yet: confirming whether a slowdown was real.

“Amodei has an extremely naive interpretation of what needs to be done there,” says Maurice Chiodo, a mathematician at the University of Cambridge’s Center for the Study of Existential Risk, who says he has audited around 30 AI companies. One immediate concern, he says, is what exactly an auditor needs access to. That could include model weights and the numerical parameters that have been shaped during training, as well as internal evaluations, red-team transcripts and incident reports. Chiodo argues that data aren’t enough. “Giving an auditor access to nothing and giving them access to a million documents has exactly the same effect, which is: they can’t get anything done,” he says. “These auditors need access to people, primarily.”


On supporting science journalism

If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.


Yet embedding auditors alongside employees, as Amodei suggested, could compromise their impartiality. “The auditors get their badge and their desk, they go to staff drinks nights on Friday night, and they enjoy it,” Chiodo says. “They really become part of the company, which makes it difficult to criticize it because the staffers become almost your friends.” Lilian Edwards, a professor emerita of law, innovation and society at Newcastle University in England and director of Pangloss Consulting, calls that setup “an absolute recipe for cultural capture.”

Independence would also depend on auditors’ freedom to speak out afterward. Amodei said Anthropic would retain the right to redact security-sensitive, legally privileged or proprietary information. “If you’re writing that in your first proposal, [that you] reserve the right to redact and hold stuff back, you’ve already lost the game in terms of safety,” Chiodo says. Amodei explicitly said Anthropic could not redact a finding simply because it was unfavorable. But Edwards is still wary. “Redactions are going to be a political and commercial question, not a technical one,” she says.

Then there is the model being audited. In August researchers at Transluce, an independent nonprofit AI lab, reported that frontier models behave differently depending on whom they believe they’re talking to. The team varied the identity presented to a model while keeping the underlying tasks the same. “The model actually adapted its behavior based on who it was talking to,” says Jacob Steinhardt, a computer scientist at the University of California, Berkeley, who leads Transluce. The effects were small on average. But when conversing with AI lab employees, the model tended to give answers that were more cautious and to use reasoning that was longer and more critical. Older models sometimes revealed in their reasoning traces, the intermediate steps a system generates before answering, that they recognized the person they were talking to. “For newer models, that's actually no longer visible,” Steinhardt says. That makes the behavior harder to spot. “We don't really know how to solve it as a field.”

Testing environments cannot be trusted alone,” Chiodo says. “You need to find a way to sample the model when it doesn't think that you're looking.” Steinhardt, though, is not convinced by sampling random AI interactions for human evaluation. “It doesn't let you anticipate new problems before they happen,” he cautions.

There is also the problem of finding enough qualified people to do the work. An audit team, Chiodo argues, needs expertise that mirrors the development team role for role. “If there's a development role that's not reflected in the audit team, then that role can't be audited,” he says. Those experts can make far more working for the AI companies themselves. “I've been shouted at, sworn at, cursed at, and told I can't speak to developers anymore,” Chiodo says. “No one's going to clap for you.”

California is already trying to formalize a broader AI-auditing ecosystem. On September 9, Governor Gavin Newsom signed AB 1405, which directs the state to create an AI Auditor Registry by January 1, 2029; beginning then, only registered auditors may conduct certain audits required for compliance with state law. A companion law, SB 813, tasks the state with deciding who qualifies to independently vet AI systems. Neither law requires frontier developers to submit to an audit to build or deploy a model.

Edwards doubts that voluntary oversight will overcome the commercial incentives to release their latest, most powerful models. “Efforts would be made to fudge it,” she says. “I cannot personally see voluntary certification making a difference between existential AI risk and evading it.”

Anthropic says it intends to bring an embedded external review team inside the company “in the near future.” Chiodo doubts independent auditing can be built fast enough to keep pace with frontier AI development. “Amodei's suggestion of external auditors, that's years into the future,” he says. The next generation of Claude is unlikely to wait that long.

Subscribe to Support Independent Journalism

Great science journalism requires human expertise, time, effort and creativity. And it costs money. That’s why I and the journalists here at Scientific American hope you’ll join our community.

When you subscribe, you are supporting staff and freelance journalists who are passionate about telling science stories that are true, important and compelling. Our editors and reporters are often experts in their fields, which means they understand the nuances of big discoveries and can untangle the breakthroughs from the hype. With a subscription, you are also supporting rigorous fact-checking to ensure the words we publish are precise and accurate. And you’re supporting original illustrations, graphics and photos that bring you closer to an advanced laboratory, an ice sheet in Antarctica or a space mission in orbit. You’re helping us craft other types of high-quality journalism as well: Our newsletters are carefully written, edited and curated by staffers you have or will come to know and love. Our Science Quickly podcast is based on original reporting, collaboration with editors and scientists and exacting production.

Subscriptions keep this engine running so we can continue to deliver thoughtful, rigorous and independent science journalism to you. In an era of viral misinformation, this work is crucial. If you value what we do, I hope you’ll consider joining us as a subscriber

Thank you,

Jeanna Bryner, Editor in Chief, Scientific American

Subscribe