On the twenty‑fourth of July, Anthropic released Claude Opus 5. The launch post is a good one, and it is entirely about economics: within half a per cent of Fable 5’s peak score at half the cost per task, a new effort dial, cheaper long‑horizon work. It mentions “much stronger visual outputs” once, in passing. It does not use the word game. Not once.
By the following evening the internet had held a vote and decided that games were the point.
Somebody described a snowboarding game and got one back, working, on the first pass. Somebody else described a first‑person shooter and got fifty‑five thousand lines of code with no art assets in it at all, because the model generated the textures, meshes, animations and sound procedurally, from code, at load time. Then a physics demo where snow remembers every footstep you leave in it. A Minecraft clone with fifteen biomes. A submarine game where the model chose a sixteen‑colour palette and hand‑dithered the textures to fake gradients, which is not a technical decision. That is art direction.
None of these people made a finished game before breakfast, whatever the headlines suggest. I’ll come to that, because it matters. But something did change, and it is worth being precise about what.
What people actually built
These are the ones worth your time, in their makers’ own posts. A lot of this work has been going round the feeds this fortnight with the credits filed off, so it seems only right to send you to the source.
A snowboarding game, described once and running on the first pass.
— Alex Ermolov (@alex_erm) view the original post
A first‑person shooter where nothing you can see is an asset file.
— Matt Shumer (@mattshumer_) view the original post
A submarine game rendered in sixteen dithered colours.
— Pietro Schirano (@skirano) view the original post
The prompt and the full source, both published, so you can study the input rather than admire the output.
— Meng To (@MengTo) view the original post
The words “one‑shot” are doing a lot of work
Here is the honest version, and I think it makes the story better rather than worse.
Matt Shumer’s shooter is the most‑shared artefact of the whole wave, and to his credit he published the entire repository: the seed prompt, the architecture, the critic logs, all of it. Open it up and “one prompt” turns out to mean one human prompt driving a fleet of orchestrated sub‑agents, iterating for hours, with separate critic agents checking each other’s work. The seed prompt itself contains the instruction to loop on each item and have a sub‑agent inspect it visually.
And then the part almost nobody quotes. The stated goal was to match a modern Call of Duty. The README answers its own question in two words: It does not. The critics scored the result 5.05 out of 10. In a blind A/B against real Call of Duty frames, every critic in every round picked the real one.
I find that genuinely more impressive than the marketing version. A person with an idea and no art team got to a coherent, playable, procedurally‑generated 3D shooter in a day, and was straight enough to publish the score that says it is a five out of ten. Five out of ten did not exist at this price a year ago. Five out of ten is a prototype you can put in front of someone.
The receipt that isn’t a vibe
Viral clips are not evidence. There are no standard tasks, no rating scale, and no control over how many attempts a model gets before somebody posts the good one. So here is the measured version.
ARC‑AGI‑3 is a benchmark built out of games the model has never seen, specifically to test whether a system can work out novel rules by playing rather than by remembering. Opus 5 scored 30.2 per cent on it, roughly four times the previous best, and the ARC Prize team verified the run themselves. Worth the caveat: that was at high effort, and on a different interactive benchmark Opus 5 ties its rivals rather than pulling away.
So: not a solved problem. But not a vibe either.
What this actually changes
I have spent a career around the moment where an idea has to survive contact with production. Somebody has an instinct, and then a team spends six weeks finding out whether it was any good, and by the time they know, the instinct has been sanded flat by the process of testing it.
That gap is where most good ideas die. Not in the pitch, not in the execution, but in the expensive middle where you cannot afford to find out.
What happened over the last fortnight is that the middle got cheap. Not free, not finished, not a substitute for the people who can actually make things — a five out of ten is still a five out of ten. But the cost of finding out whether an idea is worth pursuing has dropped by something like two orders of magnitude, and it dropped for everybody at once.
If you have been sitting on an idea for a thing that should exist, this is the fortnight the excuse expired. Describe it properly, let it run, and look at what comes back. It will be a five out of ten.
Five out of ten is a beginning. It never used to be available on a Tuesday.
Si