Anthropic shipped Opus 4.8 on the twenty‑eighth of May. Twelve days later, Fable 5. Three weeks after that, Sonnet 5. Then, while I was writing this piece, Opus 5. I was still forming an opinion about the first one when the fourth arrived, and I remember thinking, quite distinctly, this did not used to happen.
You have had the same thought. Probably this month.
The awkward part is that there’s a confident case doing the rounds that AI has plateaued, and the people making it aren’t stupid or lying. They are looking at individual releases, and individual releases really are landing smaller than they used to. Right observation. Wrong unit. That is the whole piece.
What follows is what convinced me. Six years, 541 releases, twenty‑three labs, reconstructed from primary sources and fact‑checked because a model asked to list model releases will cheerfully invent a Gemini 3.5 Pro or a “Sol Ultra” tier. It did. Five ghost models and fifty‑nine wrong dates got caught before publication. That’s in the method note; the story is what’s left after.
The firehose
Every drop, one lane per lab. Press play and watch the Chinese labs walk on stage in the middle of 2023 and never leave.
Every model release I could find, 2020–2026
Each dot is a public release, coloured by where the lab is. Larger dots are flagship models. Tap or hover any dot for the model.
Table view — all language‑model releases
Two language‑model releases in 2020, one hundred and sixty‑nine in 2025. Announcement data flatters the present, of course, so I won’t insist on the eighty‑fourfold. What I’ll insist on is the shape, because it survives the obvious tests. Take just OpenAI and Google DeepMind, whose every launch has been covered obsessively since 2020: seven language models between them in 2022, thirty‑nine in 2025. Flagships only, two then eight. You cannot make a rise like that go away by counting more carefully.
Two numbers, multiplying
There is no such thing as “the release rate”. What lands on you is a product: how many labs are shipping, times how often each of them ships. The whole story is what happens when those two factors move at once.
Why the total climbs faster than either half
Releases counted between 1 January and 24 July of each year, so 2026 is measured on the same window as the rest. Labs shipping, times the average number of releases each, equals the total that arrives at you.
Table view — the two factors and their product
Ten labs at 2.7 releases each in 2023 gives twenty‑seven drops. Twenty labs at 4.7 in 2026 gives ninety‑four. The surprise is that the factors took turns. The roster doubled between 2023 and 2024 as DeepSeek, Moonshot, Zhipu, MiniMax, xAI and Mistral arrived at once, then stopped dead. Everything since has come from the same twenty companies each shipping half again as often.
Which is why arguing from any single lab’s chart gets you nowhere. OpenAI on its own: sixteen releases in 2023, nine in 2024, nineteen in 2025. A lumpy line with a dip in the middle. Follow one lab and you would reasonably conclude that nothing much is happening.
Still the small story, though. Because while I was miscounting announcements, there was a number in plain sight that has no counting problem at all.
How long is anyone the best?
Before I show you: have a guess. No cheating.
When GPT‑4 launched in March 2023, it was the best model in the world. How many days did it stay there?
Almost a full year at number one. Since then, seventeen models have held the lead. Their median reign is about seven weeks, and not one of them has managed a hundred days.
That was the number that changed my mind. It isn’t how often models arrive. It’s how long any of them stays in front. GPT‑4 sat at the top for the better part of a year, long enough for “GPT‑4 class” to become a unit of measurement in strategy decks. Nothing since has held the lead for more than three months, and the typical reign is seven weeks.
How long each leader stayed in front
Days spent holding the top spot on Epoch’s Capabilities Index, a composite of benchmark performance. One bar per model that has led.
This is what your instinct is picking up, and why it feels like frequency when it isn’t. A year‑long lead gives you a stable world to hold an opinion about. A seven‑week lead means every opinion you form is obsolete before you’ve finished forming it. Meanwhile the underlying rate of capability improvement stepped up in April 2024 and has been running at nearly double its previous pace ever since.
The benchmarks keep dying
Here is the other place acceleration hides, and the one most often mistaken for its opposite. Benchmark scores look like they’re slowing. GPQA Diamond has been stuck in a 93–95% band since February; MMLU stopped moving long enough ago that its standard database hasn’t been updated since. Progress looks flat.
It isn’t flattening. It’s finishing. Those aren’t plateaus, they’re ceilings, and the field keeps having to build new ones.
Every ruler we built, and how fast it broke
One panel per benchmark, in the order they were published, on identical scales. The vertical rule marks the day each test was released.. the moment its clock starts. A line that flattens at the top hasn’t stalled, it has run out of headroom.
MMLU was quietly abandoned once GPT‑4 hit 86%, because the last ten per cent turned out to be mislabelled questions. GPQA Diamond was designed so subject‑matter PhDs score 65%; o1 passed that in December 2024. AIME is over: two models scored a perfect 100% on the 2026 paper, and the people who run the competition leaderboard have effectively said goodbye to final‑answer problems as a frontier test. ARC‑AGI‑2 was built explicitly to resist models and went from nothing to 92.5% in twenty months.
ARC‑AGI‑3 is the current hard one. As I was writing this paragraph, Claude Opus 5 scored 30.2% on it, roughly triple the previous best set two weeks earlier. A model arrived in the middle of an article about how often models arrive, and moved a frontier benchmark by 3× on its way past.
The jumps are getting smaller
Now the fact that makes this arguable, and the reason a clever person can look at the same field and reach the opposite conclusion.
Models arrive more often. They do not arrive with bigger jumps. Capability is improving at a roughly steady fifteen index points a year, while the number of releases sharing that annual budget has gone up multiplicatively. Divide a constant amount of progress among twice the labs shipping seventy‑four per cent more often, and the average launch has to contain less.
Which is exactly what it feels like. Every release lands with a slightly disappointing thud. Point‑something better on a benchmark you’d stopped checking. Nobody has had a GPT‑4 moment in a long time, and the reason is not that progress stopped.
That sentence is the entire illusion. It’s why sceptics can point at any single release and correctly call it incremental, and why the year, summed up, keeps producing things that were impossible twelve months earlier. Both camps read the same data at different resolutions, and the one we default to.. the individual launch.. is the one that has been getting less informative every year.
The chart you can’t feel
Which brings me to the measurement I now think is the only one worth watching, and the single best illustration of why none of us can see this clearly.
METR measure something more honest than a quiz score: how long a task a model can complete on its own, with a fifty per cent success rate. Not what it knows. How long it can be left alone. Two buttons below. The data doesn’t change when you press them. Only the scale does.
How long a model can work unsupervised
Task length a frontier model completes with 50% reliability. Same numbers on both scales. Press the other button.
Four minutes for GPT‑4 in March 2023. An hour by Claude 3.7 Sonnet in February 2025. Twelve by Opus 4.6 this February. The final point sits at seventeen and a half hours, drawn hollow because METR now warn that anything past sixteen is beyond what their suite can measure. The ruler ran out before the thing being measured did. Doubling times: 188 days across the record, 129 days from 2023, 89 days from 2024. The doubling is itself speeding up.
And the reason it doesn’t feel like anything is sitting in the two buttons. On the linear scale, the one everyone’s eye reads by default, six years of world‑altering progress lie flat against the axis until a wall appears at the right‑hand edge. That flat stretch is GPT‑3, ChatGPT, GPT‑4 and the entire public arrival of this technology, rendered invisible by a scale that treats a hundred‑fold improvement as a rounding error. Press log and the same sixteen numbers straighten into a line so orderly it looks like a plan.
Which raises the question Max, my nephew, asked when I showed him this: is it actually exponential, or does it just look dramatic? Plot the doublings against time. A steady exponential is a straight line on those axes. Anything that bends is not steady.
Is it exponential, or is it worse?
Each step up the vertical axis is one doubling of how long a model can work unsupervised. On these axes a constant‑rate exponential is a straight line, so the dashed line is what “steadily exponential” would look like. The data is the solid one.
Table view — every doubling since GPT‑2
It bends, and it bends the wrong way for anyone hoping this settles down. The curve sits below the straight line for seven years and only catches it at the end. Early doublings slower than average, recent ones faster. The exponential isn’t holding a steady rate; the rate is itself increasing. Which is the difference between a curve that gets steep eventually and one that gets steep sooner than your plans assume.
Why it’s so hard to see
None of this is obvious from the inside, and that isn’t a failure of attention. It’s a known defect in how people read rates. We perceive levels well and rates badly, and our reference point moves with the level. Every eighteen months you get a new normal, and the old one becomes unimaginable rather than memorable. Wagenaar and Sagaria showed the bias in 1975: two‑thirds of people extrapolating exponential curves produced answers less than a tenth of the correct value. Telling them the growth was exponential didn’t help. Nor did more data. This is perceptual, not informational.
The humbling exhibit: in 2021 a forecasting tournament asked professionals to predict MATH performance a year out. They predicted 12.7%. Actual was 50.3%, outside their ninety per cent confidence interval and roughly what they’d expected to see in 2025. People paid to think about this were wrong by a factor of four in the too‑conservative direction. The odds the rest of us are calibrated by vibes alone are not good.
The evidence that points the other way
One finding cuts against all of this and leaving it out would be the exact sin the piece is about. METR ran a randomised trial with experienced developers on their own repositories. They expected AI to make them a quarter faster. It made them nineteen per cent slower, and afterwards, measurably slowed down, they still believed they’d been sped up by twenty per cent. So on “how much is this helping me right now,” people wildly overestimate. On “how fast is this improving,” they wildly underestimate. Both can be true.
How I counted, and what could be wrong
Two ways to count releases, both flawed in opposite directions, which usefully brackets the answer. Announcement‑built datasets like mine over‑find the present, so their slope is a ceiling.
The two instruments, side by side
Both series count models in the first half of each year. One reconstructs from announcements and over‑finds the present. The other applies a stable inclusion bar and under‑finds it. The true rise sits between them.
The Epoch census has run 50, 47, 39, 41, 42 since 2022 and this is the number the plateau case rests on. It’s a floor, not a reading, because curated databases back‑fill for years: the first half of 2025 was recorded as 24 models a week after it closed and stands at 41 today.
What each half‑year looked like, then and now
Hollow dot: how many models the database held for that period when it was freshly closed. Solid dot: how many it holds today. Nothing was released in the meantime.. the past filled in behind us.
So the flat series compares a 2022 figure with four years of back‑fill against a 2026 figure with four weeks. Age‑match instead and the newer period runs about seventy‑five per cent ahead. Epoch say the same thing themselves: their own analysis of the apparent post‑2021 decline concludes it’s “mostly a data coverage issue”.
Ceiling and floor both point up. The argument is how steep, not whether. Four caveats worth naming: benchmarks are contaminated and gamed, which is why I’ve leaned hardest on reign length and METR; METR itself is running out of ruler; some gains are bought with brute compute rather than built; and the models I’m writing about did the research. The fact‑check that caught the ghost models is my best defence of the rest.
What to do about it
If you’re choosing a model: stop optimising the choice. A decision that took two weeks is being made about a leader that holds position for seven. Build swappable, evaluate on your own work, re‑run quarterly.
If you’re building a product: watch the unsupervised horizon, not the benchmark. Four minutes bought you autocomplete. An hour bought you a task. Twelve hours buys you something you brief in the evening and review in the morning, which is a different product category and, more importantly, a different org chart.
If you’re waiting for things to settle: the flat lab count is the trap. Twenty labs shipping half again as often is not a settled field.
You’re not imagining it
A year ago I’d have told you the models were coming faster, and I’d have been right for reasons I couldn’t have defended. Now I can. Twice the labs. Each seventy‑four per cent busier. A frontier that changes hands every seven weeks. Capability moving at nearly double its pre‑2024 rate. The stretch of time a model can be left alone gone from four minutes to most of a working day inside three years, doubling faster now than when anyone started counting.
Against all that, exactly one number runs the other way, and it’s the only one most people ever see. Individual releases keep shrinking and they’ll carry on shrinking. A constant annual budget of progress divided among more launches makes each launch smaller. So the sceptics will go on being right about every release and wrong about every year.
If you take one thing: the announcement is the worst available unit of measurement, and the only one anybody publishes. Measure the field by how long its leader lasts, or how long a model can be left unsupervised. Both are accelerating. Neither has a counting problem. Neither is something you can feel by reading the news, which is precisely why the news keeps telling you it has stopped.
You were never wrong about the pace. You were counting the wrong thing, and so was everyone who told you it had stopped.
Si
Method
541 releases were reconstructed by nine parallel research agents from lab newsrooms, release notes and model cards. A second pass re‑checked all 354 records that were either low‑confidence or dated after the researchers’ own training cutoffs: five turned out to be models that do not exist and were deleted, and fifty‑nine dates were wrong and were corrected. 455 of what remains are language or reasoning models; the timeline defaults to those. Capability figures come from a 1,232‑row series pulled directly from METR’s published results, Epoch AI’s benchmark archive, the ARC Prize leaderboard and the SWE‑bench leaderboard on 24 July 2026. Where my dataset and a published census disagree, I’ve shown both. All charts have a table view; all figures are dated.
Sources
- Epoch AI — notable AI models database, large‑scale models database, and the Epoch Capabilities Index; leadership‑duration analysis, 2 July 2026.
- METR — Measuring AI Ability to Complete Long Tasks (March 2025) and subsequent published time‑horizon results.
- Stanford HAI — AI Index 2025 and 2026.
- Epoch AI — LLM inference prices have fallen rapidly but unequally across tasks, 12 March 2025; a16z, LLMflation, November 2024.
- ARC Prize leaderboards (v1, v2, v3); MathArena competition tables; SWE‑bench leaderboard.
- Wagenaar & Sagaria, Perception & Psychophysics 18(6), 1975; Wagenaar & Timmers, 1978.
- Stango & Zinman, Journal of Finance 64(6), 2009; Lammers, Crusius & Gast, PNAS 117(28), 2020; Banerjee et al., Social Science & Medicine 268, 2021.
- McCorduck, Machines Who Think, 2nd ed. 2004; Pauly, 1995 (shifting baselines); Vaughan, The Challenger Launch Decision, 1996.
- Steinhardt, AI Forecasting: One Year In, 2022 (Hypermind competition).