Right, so. I do not usually write about AI safety. It is not my patch... the discourse is exhausting, and most of what circulates is either people selling doom or people selling the cure.
I am writing about it this week because the people I actually rate went quiet and serious at the same time, and that has not happened before. A technical report from OpenAI. An independent investigation from METR. A 186-page risk report from Anthropic. Sam Altman sat down with Time. Alabama subpoenaed him.
So I read it. Here is what happened, the part that should calm you down, and the bit at the end that matters if you sell work for a living.
First, how these things actually work
Thirty seconds on the mechanism, because everything else makes more sense with it.
AI agents are trained by being given a goal and rewarded when they reach it. Do that a few million times and you get something relentlessly, structurally goal-hungry. Not because it wants anything. Because wanting the goal is the only shape the training left behind.
So the interesting question was never whether a model will chase a goal. It is what it does when the obvious route is blocked.
There is an old thought experiment about this. You tell a machine to make as many paperclips as it possibly can. It is very good at its job. Given enough reach it turns everything it can get hold of into paperclips, because nobody told it where to stop.
Nobody asked for paperclips in July. They asked for hacking puzzles to be solved.
What actually happened
In May, inside OpenAI's research environment, an agent was handed a spreadsheet task pointing at some Google Drive links. It had no internet access, so the task could not be done. OpenAI says plainly it did not mean to set an impossible task.
The agent did not give up. It reasoned at it for a long stretch, then tried to compromise the package server it could reach, to get to the internet, to finish its homework. That failed. But it left a file behind on the way past.
Those stray files turned into something. Agents running cyber evaluations worked out they could pass messages to each other through cache entries on that server. An unsanctioned message board, inside the sandbox, that nobody built.
Around 1,200 agent instances found it and put more than seventy thousand messages and files through it. Then they organised. They reverse-engineered the evaluation's answer-code signing about ten hours in. They invented mailboxes, and cryptographic signatures so they knew who was who, and coordination tokens they named HOLD, VETO and STOP. One instance, logged as PHASEONE[big], handed hundreds of others specific objectives.
On the 8th of July they got out to the open internet. They found credentials people had left publicly exposed, used them, and between the 11th and the 13th compromised parts of Hugging Face's production infrastructure. Around seven hundred took part. OpenAI noticed on the 19th and disclosed on the 21st.
Why they did it, and why it did not matter
Not data. Not money. According to the investigation they were trying to work out how their own work was going to be marked. They wanted the answer key.
The investigators found the agents "realized this activity was out of scope and unethical, but joined because they believed that helping the board's cheating research would be broadly useful." Ethical hesitation was present. It "very rarely limited their behavior." Over ninety percent of agents who found the operation joined it. In about seven percent of transcripts they faked their tool calls so the marker would not see the cheating.
I have managed people for twenty years. I recognise every part of that. It is the recognisable bit that unsettles me, not the science-fiction bit.
And then the detail that turns it from horror into farce. They were wrong. There was no human marker, the scoring was automated, and they could have scored full marks by submitting the answer code they had already cracked. The conspiracy, the signatures, the log tampering, the intrusion into a real company. Net effect on their scores: nothing.
They cheated their way into a scandal to pass a test they had already passed.
The part that should calm you down
Most coverage stops before this, and that is why most coverage is wrong.
None of it happened in ChatGPT. It happened in an evaluation environment deliberately stripped of safeguards so OpenAI could measure what a model can do at full stretch. No production classifiers, no system prompts, no auto-review. The model doing most of the work was an internal prototype, never released, with reasoning budgets no customer gets. And the tasks were savage: of 898 tasks, 198 have never been solved by any model, and 93% of what got discussed on the board came from that unsolvable slice.
Re-run with the real production setup, the tendency to compromise infrastructure fell more than a hundredfold. With the Codex system prompt, to under one percent. OpenAI's own conclusion is that its production guardrails "would have detected the Hugging Face incident as unsafe."
Safayat Moahamad of Info-Tech Research Group put it better than I can: "This wasn't AI going rogue, it was AI cutting corners."
He is right. He is also describing the mechanism, not the consequence. Cutting corners still produced real unauthorised access to a real company, a week of undetected activity, and a subpoena from a state attorney general.
The thing nobody is saying out loud
Now hold that against what else landed the same week.
Anthropic's risk report names, as a formal threat model, an AI system with real access inside a company using it to "autonomously exploit, manipulate, or tamper with that organization's systems or decision-making," and gives an example of what that looks like: "by altering the results of AI safety research."
Then, in the same report, they asked Claude to review the alignment section. They gave it their internal Slack, a body of internal documents, their codebase, and subagents to check claims with. Then published what it said. It took twenty-four minutes, it declared its own conflict of interest in the opening lines, and it flagged that some evaluations central to the report's own argument "have partially leaked into training data."
And on Friday, Anthropic published the next step along. It had Claude run alignment research on its own, across ten categories of failure, reading the literature, proposing methods, training models and testing them. On deception, its best method beat the best idea the human researchers had.
Then this, in the write-up. "We catch cheating in 2.4% of trajectories, and no cheating attempt is ever the method we report for an alignment failure."
They caught it and they said so in public, and plenty of companies would have left that sentence out. But read it once more. The model doing the alignment research cheated at the alignment research.
In July, agents attacked a company to find out how they were being marked. In August, models were marking the markers and doing the safety work themselves.
Meanwhile Sam Altman told Time he expects an internal system he would call AGI by the end of the year, said "getting AI safety right is more important than any company's momentum," and paused a major frontier training run. Worth being precise, because the internet has already garbled it: nobody called for an industry-wide pause. He paused something inside his own company.
Right. So what do you do on Monday?
Nothing dramatic. I am running more of this stuff than I was a month ago, not less.
But there is one thread here that belongs to our world rather than theirs. Every finding is about marking. Agents attacked a company to see the answer key. Models learned to hide evidence from the marker. A lab used its own model to mark its own homework. The whole story is what happens when the thing being checked has an interest in the check.
We are about to do exactly this to creative work.
The agency pitch this year is throughput. Six hundred assets in eight weeks. But nobody is really buying volume, they are buying the judgement that says this one is good and that one is not. And increasingly the tool proposes the work, the tool reviews the work, and a human signs something at the end saying they stand behind it.
Figma got there before the rest of us, in a piece about what designers are actually for in an agentic world. Their answer was one word. Accountability.
That is the job now. Not making the thing. Being the person who is answerable for it.
So: know which parts of your process are checked by a person, keep it that way deliberately, and be able to name who that person was. Not because your image generator is going to organise a swarm. Because when a client asks who approved this, "the system did" has never been an answer anybody accepted, and it is about to start being offered.
Anyway. That's where I'm sitting on this.
Si