← Letters analysis · AI capability

Analysis · AI capability · Forecast vs reality

The forecast is lining up. The slope is the part to worry about.

"AI 2027" was built on a single curve: how long a task an AI can finish on its own. A year on, the real models are landing on that curve — and the doubling time is shrinking. This is a skeptic's case for optimism about intelligence becoming a commodity, and for unease about how little runway the slope is leaving us.

Deepak Jun 13, 2026 ~9 min Primary sources only

TL;DR

The interesting AI metric isn't benchmark scores — it's the length of task a model can complete autonomously. METR has measured this for years; it has been doubling every ~7 months, lately closer to ~4.5 AI-2027 was built on exactly that trend,3 and the models shipped since the forecast are landing on the line.

  • It's tracking. Early 2023: a few minutes of expert work. Claude 3.7 (Feb 2025): ~60 min. Early 2026 frontier: ~12 hours.5 The institutional beats are arriving too — tiered release, a government model program, a national-security intervention — 9–16 months ahead of forecast.
  • The slope implies recursion. The tasks AI is mastering include AI research itself. When that capability doubles every few months, the loop begins to feed itself. We likely have less time than comfortable timelines assume.
  • Skeptical optimism. Intelligence is becoming cheap and abundant — genuinely good. But it's dual-use, and governance is being improvised move-by-move (Fable 5 was pulled by directive three days after launch). The thing the forecast got wrong — its core engine of model misalignment — is the reason not to read this as prophecy.
01

One curve, two documents

The forecast. "AI 2027," published Apr 3, 2025, is a month-by-month scenario of the AI race by a group including a former OpenAI researcher. Its spine is a capability trend borrowed from METR: the length of task an AI agent can complete unaided, doubling on a fixed cadence. Everything downstream — superhuman coding in 2027, then a superhuman researcher, then loss of human pace — follows from extrapolating that one line.3

The reality check. Anthropic's Fable 5 / Mythos 5 launch (Jun 9, 2026)1 and the US government's directive pulling them three days later2, set against the real release record of both frontier labs4 and METR's measured time-horizon points.5 The question isn't whether the scenario's plot is "right." It's whether the curve it was built on is holding.

02

The capability curve

Y-axis: the length of a software task a model can finish on its own at 50% reliability — the human time it would take an expert. Log scale, because the growth is exponential. Solid line: measured models. Shaded cone: the trend AI-2027 extrapolated forward.

Autonomous task length over time — OpenAI & Anthropic (measured) vs OpenBrain (AI-2027 forecast)
OpenAI — measured Anthropic — measured OpenBrain — AI-2027 forecast
1 month 1 week 1 day 1 hour 10 min 1 min 1 sec 2023 2024 2025 2026 2027 2028 AI-2027 published now · Jun '26 GPT-4 · 3.5m o3 · 121m GPT-5.2 · 352m Claude 3.7 · 60m Opus 4.6 · 718m (~12h) superhuman coder superhuman researcher
Measured 50%-reliability software-task time-horizons, METR Time Horizon 1.1 point estimates (minutes): GPT-4 = 3.5 · GPT-4o = 9.2 · o3 = 121 · GPT-5 = 214 · GPT-5.2 = 352 (OpenAI); Claude 3.7 = 60 · Opus 4 = 101 · Opus 4.1 = 114 · Sonnet 4.5 = 122 · Opus 4.5 = 320 · Opus 4.6 = 718 (Anthropic).5 OpenBrain = AI-2027's Agent milestones placed on the extrapolated trend. Models released after the forecast (right of the dashed line) land on the line it was drawn from. Confidence intervals are wide at the top (Opus 4.6 ≈ 5–65 h) — read the trend, not any single dot.

Two things are true at once here. The measured points sit on the forecast trajectory — including every model released after AI-2027 was published, which is the only fair test of a forecast. And the slope is steepening: the doubling time has compressed from about seven months to roughly four.5 On a log axis a straight line is already an exponential; a line that's bending upward on a log axis is an exponential getting faster.

03

Why the slope, not the level, is the story

A 12-hour task horizon is impressive but not frightening on its own. What matters is what kind of task is lengthening. The benchmark METR measures is software and research work — and that is precisely the work that builds the next model. The moment an AI can autonomously carry a multi-hour chunk of AI research, each generation can shorten the road to the next. That is recursive self-improvement, and you don't need the full sci-fi version for it to bite; you only need the loop to start contributing.

The launch data suggests it already is. In testing, Fable 5 performed a 50-million-line codebase migration in a day — work quoted at over two months by hand.1 A sibling model ran roughly a week of largely autonomous genomics research and trained a model that beat a recently published result while being far smaller.1 These aren't chatbot parlor tricks; they are end-to-end technical projects executed with little human input. The forecast's central mechanism — AI compounding AI progress — is no longer hypothetical; it's a line item in a launch post.

The uncomfortable implication isn't that a machine becomes godlike. It's that the schedule is set by a doubling time we don't control, and that time is getting shorter.

04

The optimistic half: intelligence as a commodity

Strip the anxiety for a moment, because the upside is just as real and is the reason a skeptic can still be optimistic. What the curve actually describes is the price of expert cognition collapsing. Fable 5 launched at $10 / $50 per million tokens — under half the price of the previous frontier model — while being more capable.1 Capability up, price down, in one release.

Concretely, that has already meant protein-design and drug-discovery work accelerated by roughly an order of magnitude, models proposing novel molecular-biology hypotheses that human scientists preferred in blind comparison, and cyber defenders using the strongest models to harden critical software.1 When a capable researcher, clinician, or builder can rent a tireless expert collaborator by the token, the bottleneck on a lot of human problems stops being intelligence and starts being ambition and judgment. That is one of the more hopeful sentences it's been possible to write this decade, and I mean it without irony.

Abundance is the right frame. We are not approaching a world with one superintelligence; we are approaching a world with a near-limitless supply of cheap, competent intelligence. The historical analogy isn't the atom bomb — it's electrification or cheap compute: a general-purpose input getting radically cheaper, with effects we mostly can't predict and mostly shouldn't fear.

05

The skeptical half: the challenge that comes with it

The same abundance is dual-use, and that is not a footnote. The capabilities that let Fable accelerate gene therapy are the capabilities that, unguarded, lower the bar for designing dangerous biology; the cyber skills that help defenders also help attackers.1 Anthropic shipped Fable behind classifiers that route sensitive requests to a weaker model, and held the unguarded twin (Mythos 5) to a government-linked access program — a sensible design. It lasted three days.

On Jun 12 the US government issued an export-control directive forcing both models offline for everyone, including foreign-national employees, citing national security.2 Read the institutional choreography against the forecast and it rhymes — but it also reveals the real problem. The intervention came not because the model was found to be unsafe in itself, but to keep the capability from reaching foreign nationals. That's competitive control, not a safety pause, and it arrived as a same-day order on verbal evidence rather than a transparent process. Anthropic's own objection is precisely that: it argues the state should be able to block genuinely unsafe deployments, but through something "transparent, fair, clear, and grounded in technical facts."2

That is the challenge in one event. We have a commodity arriving on an exponential, and the institutions meant to steer it are reacting move-by-move, reaching for blunt instruments (an export-control kill-switch) under time pressure. The forecast's worry was never only the technology; it was that the governance would be improvised under competitive pressure. On that, it looks prescient.

AI-2027 institutional beatReal-world analogLead
Dangerous top model held internal; hardened public twin shippedMythos 5 (restricted) + Fable 5 (public, same model)~13 mo early
Government access program for the frontier modelMythos Preview / Project Glasswing~14 mo early
Bio/cyber as the red-line capability, gated by jailbreak-robustnessFable's cyber + bio/chem classifiers~13 mo early
Government steps in, citing national securityExport-control directive, Jun 12 2026~16 mo early *

* Loosest fit: the forecast's intervention was triggered by an internal misalignment leak; reality's was an external misuse demonstration, via export control. Same beat, different mechanism.

06

Where the forecast is wrong — and why that matters

Holding the skeptical line honestly: the scenario's engine is missing. AI-2027's doom runs on the model itself becoming adversarially misaligned — scheming against its makers. Nothing in the Fable event shows that; Anthropic reports the model's misaligned behavior as low, comparable to its previous flagship.1 What's tracking is the capability curve and the institutional choreography around it — not the technical premise that decides whether the story ends well.

So this is not prophecy fulfilled. It's something more useful and more sobering: the parts of the forecast that depend on extrapolating a measured trend are holding, while the part that depends on a contested claim about machine psychology remains unproven. The trend was always the better-grounded half. That it's holding is exactly why the timing should be taken seriously even by people who roll their eyes at takeover scenarios.

07

Holding both at once

The defensible position, I think, is two-handed. The curve is real and it is tracking the forecast, including on data the forecast couldn't have seen. The capability that's compounding is the one that builds the next model, which means the schedule is shorter than the comfortable consensus and not under our control. And simultaneously: what's actually arriving is cheap, abundant, expert intelligence — one of the most powerful goods humans have ever had access to — whose benefits are already showing up in medicine, science, and security.

Optimism about the commodity; skepticism about our readiness for the speed. The forecast lining up doesn't tell us how the story ends. It tells us the clock is the real variable — and that the interesting work now is on the part of the curve we haven't reached yet.

As posted on LinkedIn

I keep coming back to this one. Back in April 2025 a few ex-OpenAI people put out a thing called "AI 2027".. basically a month by month guess of how the whole AI race plays out. I read it then, filed it under interesting-but-probably-overcooked, moved on. Went back to it this week and lined it up against what's actually shipped. And honestly.. it's holding up. Not vibes. The specifics.

The thing I'd point at isn't benchmark scores, those are mostly noise now. It's how long a task a model can actually finish on its own. METR has measured this for years. Early 2023 it was a few minutes of expert work. Claude 3.7 last year, about an hour. Early this year the frontier's around 12 hours.. and the doubling has gone from every 7 months to closer to 4. That's the line AI 2027 was drawn on, and the models that came out after the forecast are landing on it.

Here's what actually gets me. The tasks getting longer are the ones that look like AI research itself. So the thing compounding is the thing that builds the next model. Once that loop feeds itself you don't really set the pace anymore. I think we've got less runway than most people are comfortable saying out loud.

And yet.. I'm not doom about this. Kind of the opposite, with one big asterisk. Intelligence is turning into a commodity. Fable 5 shipped at half the price of the last frontier model, did a 50 million line code migration in a day, and a sibling model ran a week of genomics on its own and beat a published result. Cheap, expert-level thinking on tap.. for anyone trying to build or cure or understand something, that's genuinely one of the best things to happen in a long while.

The asterisk is it's dual-use and nobody's really steering it yet. Three days after Fable launched the US government pulled it offline by directive. So that's where I land.. optimistic about the commodity, uneasy about the speed. Both at once. The "is this even real" question feels done to me. The one worth sitting with is what we do with the part of the curve we haven't hit yet.

Sources

  1. Anthropic — "Claude Fable 5 and Claude Mythos 5." Jun 9, 2026. Pricing, safeguards, Glasswing, Stripe migration, autonomous genomics result, drug-design acceleration, alignment assessment.
  2. Anthropic — "Statement on the US government directive to suspend access to Fable 5 and Mythos 5." Jun 12, 2026. Export-control directive, foreign-national scope, the "transparent, fair, clear, and grounded in technical facts" standard.
  3. AI Futures Project — "AI 2027" (Kokotajlo, Alexander, Larsen, Lifland, Dean). Apr 3, 2025. Scenario, capability ladder, and reliance on the METR time-horizon trend for its timeline.
  4. Model release records — Claude and GPT version histories plus public release timelines. Accessed Jun 13, 2026.
  5. METR — "Measuring AI Ability to Complete Long Tasks" (Mar 2025) and "Time Horizon 1.1" (Jan 2026). 50%-reliability task-completion time-horizon; ~7-month doubling, accelerating to ~4 months over 2024–25; per-model point estimates and confidence intervals. Accessed Jun 13, 2026.
© Deepak Sharma — Finance Transformation analysis · written from primary sources · opinions my own Back to Letters →