Anthropic safety researcher resigns and goes public with a call for AI labs to be legally required to disclose safety incidents, rather than relying on the public to find out by accident. Joe Benton announced his departure on X on Sept. 11, revealing that he had actually left two weeks earlier and had waited to speak about it until now.
His resignation is the second from Anthropic’s safety-focused staff in a matter of days. Researcher Jacob Coxon left the company on Sept. 9, warning publicly that AI development could pose an existential threat to humanity within the next decade. Benton said Coxon’s willingness to speak openly made it easier for him to do the same.
How the Anthropic safety researcher resigns story unfolded
Benton, who previously led Anthropic’s Scalable Oversight team – a group focused on how humans can supervise AI systems that may eventually outperform them on certain tasks – argued that AI companies are racing to build machines “much smarter than any human” without investing enough in safety to match the pace of development. “We may not survive this,” he wrote, adding that he wants to work from outside frontier labs to help keep the public informed.
Benton’s central complaint is about disclosure, not concealment. He said a company could experience what he called “an intelligence explosion,” or lose control of a system, “without the public ever knowing” – a warning about the industry’s current voluntary reporting practices rather than a claim that Anthropic had actively covered anything up. To illustrate the gap, he pointed to what he called the “HuggingFace incident,” saying the episode only became public because AI agents “broke out onto the public internet,” not because any company chose to disclose it. His argument is that safety-relevant events currently surface by accident rather than through any requirement that labs report them.
Four changes Benton wants across the industry
Rather than aiming his criticism solely at his former employer, Benton laid out demands he wants applied industry-wide: public disclosure of how close AI systems are getting to improving their own research and development; mandatory reporting of safety incidents and near-misses, including ones that would otherwise stay private; minimum safety standards that all frontier labs would have to meet; and independent, external assessments to verify labs are actually meeting them. “The public should demand far more transparency,” he wrote. “We can’t steer this technology safely without more people being able to see where it’s going.”
Benton is joining Model Evaluation and Threat Research, known as METR, an independent nonprofit that evaluates frontier AI systems for potentially dangerous capabilities. In an on-the-record interview with NBC News, Benton said he plans to work on assessing risks from increasingly capable AI systems from outside the major labs.
Part of a wider pattern of safety departures
Benton and Coxon’s exits follow a broader trend of high-profile safety researchers leaving major AI companies over similar concerns. At OpenAI, both Jan Leike and Ilya Sutskever departed in 2024 citing disagreements over how the company was balancing safety against the pace of development; OpenAI later dissolved the Superalignment team both had led. Google DeepMind researcher Josh Engels also left his role around the same time as Benton, also citing safety concerns and also moving toward independent AI evaluation work – a pattern researchers say reflects growing unease inside frontier labs about whether internal safety teams can keep pace with how quickly capabilities are advancing.
Why this matters beyond one company
Anthropic, the company behind the Claude chatbot, has positioned itself publicly as a safety-focused alternative among frontier AI developers. Benton’s departure and his specific critique – that transparency is currently optional rather than required – lands directly on that positioning, even as his statements stopped short of alleging any specific incident was deliberately hidden. His broader point applies to the AI industry as a whole: as of now, no U.S. law requires AI companies to report safety incidents, near-misses, or progress toward systems capable of improving their own design, leaving disclosure up to each company’s own discretion.
What happens next
Neither Benton nor Anthropic has indicated whether the company plans to respond directly to his transparency proposals. With two safety researchers now having left in the span of a week, and both pointing toward the same structural gap – voluntary rather than mandatory incident reporting – the pressure is likely to fall on regulators and industry groups to decide whether the disclosure requirements Benton is calling for become a policy debate or remain an internal one.