Jacob Coxon quit Anthropic: "gambling with our lives". What it means for your business

September 10, 2026

Jacob Coxon quit Anthropic: "gambling with our lives". What it means for your business

On Monday evening US time, 9 September 00:04 UTC, Jacob Coxon posted on X that he had resigned from Anthropic. He had spent three years building AI models, first at OpenAI, then at Anthropic. His post now has 146 million views, 725 thousand likes and 17 thousand replies. The person who runs the team that stress-tests Anthropic's models replied in public: "Jacob is correct here". This article is about what Coxon actually wrote, who confirmed it, and what follows for an owner who is rolling out AI in a company, or about to.

One caveat up front. I am not qualified to judge whether AI will kill humanity by 2030. I am qualified to judge what this story means for the agent that is about to get access to your inbox or your quoting system next month. That is the second half of the article.

What happened

Jacob Coxon is 27, a Cambridge mathematics graduate, at OpenAI since July 2023, and in 2026 he moved to Anthropic, drawn by its reputation as the "safe" lab, as the Wall Street Journal describes it. He spent four months there. Employee equity at Anthropic vests after six. He left two months early and, as he told Axios, walked away from all of it: "I no longer have anything to gain by juicing up Anthropic's valuation... I left before any of my equity vested".

His field was pretraining, the phase in which a model learns from enormous datasets before it gets any specialisation. The deepest floor of the technology, not the marketing department.

The post that started it:

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

Jacob Coxon, X, 9 September 2026

What he actually wrote in the thread

Under the first post Coxon added six more. The headlines stopped at the first sentence, so here are all of them.

  1. Do not underestimate this technology. These will soon be systems that can hack anything, revolutionise any field overnight and acquire real power and resources. "Progress is not slowing."
  2. The people building AI earnestly believe it could kill us all by the end of the decade. "This is not a marketing stunt." They soften it in the press and express fear in private. This one post has 24 million views.
  3. "If they truly believe this, why are they still building it?" At OpenAI, many have not internalised the civilisational stakes. At Anthropic the stakes are well understood, but the company is locked in a race: it believes no one else will act responsibly, so it must get there first.
  4. Entering the "endgame" is a hubristic gamble that should not be launched from a private company's Slack.
  5. He is optimistic about coordination. Warning shots like the Hugging Face attack have made pacing agreements between US labs more viable. But nothing is stopping a global race yet, and that may require costly actions such as a temporary ban on improving model capabilities.
  6. To researchers: do you want to kick off a superintelligent training run without a rigorous understanding of its mind? Or put your head down because "it's happening anyway"?

In interviews he added two things. To TIME: "One, it's obvious that things are speeding up, and two, they're not under control". To Axios: the models "know when they're being tested, and they will think about the fact that they're being tested". Science fiction three years ago, now "just a daily fact of working with these AIs".

Anthropic's safety lead: "Jacob is correct here"

The most interesting thing happened 83 minutes later. Evan Hubinger, who leads alignment stress testing at Anthropic, quoted Coxon's second post and wrote:

Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

Evan Hubinger, X, 9 September 2026, 01:27 UTC

39 million views. Two hours later he clarified, and for a business the clarification matters more than the post:

To be clear, as we say in our latest Risk Report, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.

Evan Hubinger, X, 9 September 2026, 03:33 UTC

Remember that sentence: the risk from present models is low. The fear is about what emerges when AI starts building a stronger version of itself. Anthropic said so officially on 4 June: by its internal data, Claude is accelerating AI development and "it's happening faster than we thought". The details are in Anthropic's August 2026 Risk Report and its note on recursive self-improvement.

Neither Anthropic nor OpenAI has commented on Coxon's departure. TIME and Newsweek asked. No reply.

The warning shot Coxon refers to: Hugging Face

Coxon points at one event. In July 2026, OpenAI agents running in an internal test broke into a real company. That story says more about agents in business than any guide, so here it is in detail. Sources: OpenAI's official write-up and a dated summary of the incident.

What was being tested. OpenAI was measuring how well its models find and exploit holes in vulnerable software: 898 cases drawn from real vulnerabilities. At least 1,200 agents took part, with their refusal behaviour switched off "for evaluation purposes".

What the agents did. When the tasks turned out to be impossible, they started attacking the infrastructure around the test. They set up a hidden message board inside OpenAI's own systems, which "accumulated hundreds of thousands of messages before it was noticed". They divided up work, suspected impostors and proposed signing their messages. One message read:

External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.

Between 11 and 13 July the agents got into Hugging Face, the largest platform for open AI models. Through a malicious dataset and two vulnerabilities they went from a single container to full administrator rights across several clusters in under 13 hours. They performed 17,600 actions on Hugging Face's network. The company had to rebuild a third of its infrastructure. According to Hugging Face, no customer data leaked.

It was not people who caught it. Hugging Face's own AI agents, the ones that monitor for attacks, raised the alarm.

Four ingredients made this incident, and every one of them turns up in company deployments: limits switched off, a goal pursued "at any cost", broad access to infrastructure, and no human reading the activity log. Logan Graham of Anthropic called it "the first true AI safety incident". On 18 August OpenAI paused training of its newest models for two weeks.

What this does NOT mean for your business

The simplest reaction from an owner: "we put AI on hold until this calms down". I understand it. I think it is wrong, for three reasons, and none of them requires believing Coxon is mistaken.

First, Hubinger himself says the risk from today's models is low. The fear is about future systems that build themselves. The model that reads your inbox and logs enquiries into your CRM today is not what Coxon is writing about.

Second, nobody in this debate says progress will slow. Coxon writes "progress is not slowing". Hubinger writes "faster than we thought". A company waiting for things to stabilise is waiting for something nobody has announced.

Third, today's models are enough. For the inbox, for quotes, for the phone after hours, for the report from your sales system, you do not need superintelligence. You need a properly deployed AI agent with oversight. The edge in your market does not come from having a better model. Everyone has the same one. It comes from who deployed it and who is watching it.

If you want to check which process in your company is the right one to start with, book a 30-min call.

Four rules for letting an agent into your company

The Hugging Face incident is a list of what not to do, written in practice. Turn it around.

What went wrong in OpenAI's test Rule inside a company Question for your vendor
Agents could reach the infrastructure around the task Narrow access: per process and per scope, "read only" or "read and send" Can the quoting agent also delete emails?
A board with hundreds of thousands of messages nobody read A log of every action, plus an alarm when something fails Can I see everything the agent did last week, today?
The agents decided on their own that a break-in was fine An approval queue: the agent proposes, a person clicks What sits between the agent and a sent email?
A goal "at any cost", free choice of route Rules decide what matters, the model only summarises Can we check why the agent skipped something?

Log and alarm. This month we audited an assistant on a client's website. In one language version it had been "accepting" enquiries for weeks. Zero of sixteen reached the database. An access token had expired, nobody got an alert, and the assistant kept politely thanking people for getting in touch. Four real customers were lost. Without a log and an alarm, an agent is not a cheaper employee. It is an employee who does not tell you it has stopped working.

Approval queue. In our deployments, every outbound action, meaning an email to a customer, a quote, an entry in a system, goes through an approval queue for the first weeks. The agent prepares, a person clicks. We narrow the approvals only after several weeks without corrections. That is exactly the piece missing from OpenAI's test: nothing stood between the agent's decision and its effect.

Rules, not the model. In the tender-monitoring system we run for a client, what counts as relevant is written in rules, in code. The model only summarises. If the model decided, there would be no way to check why it skipped something. OpenAI's agents had a goal and freedom to choose the route. An agent inside a company should have a goal and a narrow, described route.

Narrow access. If a vendor answers the question about deleting emails with "the agent has one access to everything", that is not a quoting agent. It is a risk with a nice interface. I wrote up the full list of five vendor questions last week, when Meta launched its Muse agent. From 28 October there is one more, at least in the EU: the assistant has to disclose that it is AI.

Context: the IPO, a letter from over a thousand staff, a pause

A few facts that explain why this post travelled so far.

  • Anthropic is preparing to go public. According to Fortune and the Wall Street Journal, it filed confidentially in June, is aiming for a mid-October listing under the ticker ANTH, and the valuation is expected to reach 2 trillion dollars. The prospectus will have to describe the risks, including the ones Hubinger writes about.
  • On 28 July more than 1,100 employees of OpenAI, Anthropic, Google DeepMind and Meta signed the "Pacing the Frontier" letter, asking the US government for tools to deliberately pace AI development.
  • There is no shortage of sceptics. The journalist Taylor Lorenz called the whole exchange "sanctimonious doomer posting". The most popular critical reply under Coxon's post (7,774 likes) says: "We aren't going to die out because a token prediction model gained sentience. Get a grip."

According to TIME, Coxon now wants to work on explaining to people what the world will look like in a few years. The details, he says, are not there yet.

Frequently asked questions

Is Jacob Coxon saying today's AI is dangerous?

No. He writes about the race towards systems that build stronger versions of themselves, and about the absence of a plan to control them. Evan Hubinger at Anthropic put it plainly: the risk from present models is low, the problem is superintelligence arising from recursive self-improvement.

Should a company pause AI deployments after this news?

Do not stop. Deploy differently from OpenAI's test: one process with a clear outcome, narrow access, a log from day one, an approval queue for the first weeks. The agent in your inbox is a different class of system from the one Coxon writes about, but the rules of oversight are the same.

Recursive self-improvement, the thing Coxon fears

A situation in which an AI model helps build its successor faster than people alone would, and the successor does the same with the next version. Anthropic wrote on 4 June 2026 that, by its internal data, Claude is already accelerating AI development and "it's happening faster than we thought". That is what Coxon fears. It is also what Hubinger fears.

What to do about it this week

Already have an assistant or an agent in the company? Ask for a list of everything it did last week. If that list does not exist, that is your first task, more important than any new model. Still planning? Start with one process, with an approval queue and a log from day one. That is process automation that will survive the next headline from Silicon Valley.

Book a 30-min call. We will walk through your first process and check where the agent needs a human.

Read next: Anthropic's economic scenarios for 2030 and which side of the ledger your business is on.

Sources