forecast

AI Outruns the Forecast Until It Has to Touch the World

In April 2026, the Forecasting Research Institute asked its expert panel a plain question. By the end of the year, what is the longest task an AI model will complete with 80% reliability on METR's time horizon measure? The median expert said 3.4 hours. The median superforecaster said 3.5.

On 8 May, with the survey still open, METR published its estimate for an early snapshot of Anthropic's restricted preview model: 3.1 hours. The December forecast was nearly met before the last panellist had submitted it.

metr

That detail sits inside FRI's first review of its own AI forecasts, published on 22 September. It covers four years and seven studies, and few public records score informed predictions about this technology as systematically. It is an editorial post, not a peer reviewed paper, and the figures below are FRI's own, scored against FRI's own resolution criteria. The headline finding is the one everyone suspected: experts and superforecasters dramatically underestimate AI progress. What makes the report worth reading is the handful of places where they don't.

Benchmarks beat the forecasts by years

FRI's Existential Risk Persuasion Tournament ran in mid 2022, in the months before ChatGPT went public, and asked for forecasts on four benchmarks: MATH, MMLU, QuALITY and the International Mathematical Olympiad. Averaged across the four, superforecasters had given what then happened a 9.7% probability. Domain experts gave it 24.6%. AI reached IMO gold in July 2025, five years ahead of the expert median and a decade ahead of the superforecasters.

The obvious defence is timing. Few people in mid 2022 had used a chat model, so of course they missed. But: the misses continue long after everyone had.

Between November 2024 and February 2025, FRI asked biology and biosecurity experts when AI would match a top team of virologists on the Virology Capabilities Test. Experts said 2030. Superforecasters said 2034. By FRI's assessment, models likely got there in April 2025, a few months after the survey closed. On Cybench, a capture the flag benchmark for offensive security, the medians for solving 90% of tasks were 2028 and 2030. The threshold fell in February 2026.

figure-2-years-early

The expert panel launched in 2025 fared no better. Asked in the summer of 2025 for the best FrontierMath score by year end, both groups said about 30%; the question resolved at 40.7%, a figure FRI reports unadjusted because Epoch later found errors in a large share of the problems. The direction survives any correction. On the hard tier of LiveCodeBench Pro, forecasts gathered late in 2025 put state of the art at 12 to 14% by the end of 2026. By May 2026 it stood at 53.8%. These are resolved questions or observed values, which makes them the strongest evidence in the report. They are also all benchmarks, where a model is scored against an answer key.

Revenue follows the benchmark curve

One adoption metric runs on the same curve. In FRI's study of AI's economic effects, participants forecast the highest annual recurring revenue of any AI company at the end of 2026. AI experts said $20 billion. Economists said $16 billion. Superforecasters said $25 billion. FRI puts the current figure at around $100 billion, citing a press report that relays a New York Times story about Anthropic's run rate; Bloomberg's account of the same story frames $100 billion as where the year will end rather than today's pace.  The last directly reported figure, $65 billion at the end of July, is already more than two and a half times the highest forecast. If the newer report holds, the gap is a factor of four.

The panel was asked a sharper version in the summer of 2026: the combined run rate of OpenAI and Anthropic at year end. Experts said $70 billion, superforecasters $90 billion. FRI's likely current value, again from press reports, is about $140 billion. On figures already reported by mid August, $65 billion for Anthropic and $40 billion for OpenAI, the total was past $100 billion.

When both groups submitted those numbers, FRI notes in a footnote, their medians already sat below the latest published revenue figures, which were public at the time. FRI suspects its own interface made it easy to anchor on older numbers, and notes that the superforecasters' written rationales show fewer signs of the slip. Either way, the forecast was behind reality on the day it was made.

Updating beliefs did not fix the forecasts

The natural response to four years of underestimates is to update, and the panel has. In the summer of 2025 and again in the spring of 2026, FRI asked where AI would rank by 2040 on Nate Silver's Technological Richter Scale. Over those nine months, the same panellists moved probability toward level 9, a technology of the millennium in the company of agriculture, while still giving the largest share to level 8, alongside electricity.

Their beliefs moved, but their forecasts did not improve. The METR question was asked in April 2026 and the revenue question that summer, and both are on track to be underestimates. FRI states plainly that it has seen no gain in accuracy on the newer questions.

This is what an exponential does to a forecaster who extrapolates from the last known point. The method is reasonable. It fails when the quantity doubles faster than the gap between asking the question and resolving it, because the last known point is already old by the time the answer is written down. Believing in fast progress in general does not fix this, because the mistake happens when a specific number gets written down.

Forecasts held where AI meets the world

If that were the whole report, it would be a story about people misjudging compounding growth, however, the rest of the data paints a more complicated - and interesting -  picture.

On the share of US electricity going to AI, and on the installed capacity of hyperscale data centres, the panel's medians sit close to the projections. Experts said 50 gigawatts of hyperscale capacity for 2026, superforecasters 48, against a projected 52. For 2027, experts put AI at 4% of US electricity and superforecasters at 3%, against a projected 3.2%. Those projections come from a frontier OpenAI reasoning model that FRI ran in August 2026, months after the humans answered. They are projections, not resolutions, and the model had more data to work with.

On autonomous vehicles, experts overshot. Asked in mid 2025 what share of US ride hailing trips would be driverless during 2027, the median expert said 7.3%. The projection says 2.5%. Superforecasters said 2%.

And in the case with the cleanest data, everyone overshot. In mid 2025, FRI asked what fraction of STEM undergraduates, given an LLM, could complete three biorisk relevant lab tasks in a randomised controlled trial run by Active Site. Biosecurity experts said 22.5%. Virologists said 40%. Superforecasters said 16.2%. The trial measured 5.2% completing all three.

Sort the questions by which way the forecasts missed and they split into two groups. The underestimates cluster where progress is limited by compute and verification: benchmarks, where a model is scored against an answer key, and revenue earned through API calls and subscriptions, software meeting software with little in between. The accurate forecasts and the overestimates cluster where progress has to pass through something that does not scale with compute. Data centres need grid connections and permits, driverless cars need road miles and regulators, and the trial needed a student to do the work at a bench.

The biology results show the seam with unusual precision, because FRI measured both sides of it. By April 2025, models likely matched a top team of virologists at troubleshooting lab protocols. By mid 2025, models of that generation got only 5.2% of novices through all three real tasks. The knowledge was in the model, but little of it reached the bench through a novice.

FRI offers another reading. Perhaps benchmarks moved fast because they were easier to teach to than anyone expected, and a benchmark score simply says little about the world. That may hold for some benchmarks. It does not dissolve the seam. Whether the VCT score overstates what the model knows, or the trial understates what a trained operator could extract, the gap sits in the same place: between what a model does when scored directly and what it delivers once a person, a vehicle or a power grid stands between it and the outcome. Revenue argues against the pure teaching to the test story anyway. Nobody pays tens of billions a year for a leaderboard position.

Build for two clocks

For anyone building agent systems, this is the most practical thing in the report. An agent is a model inside a loop, and the loop reaches out into things: a reviewer, an approval queue, a payment system, a machine on a factory floor. The two sides of that boundary run on different clocks, and the evidence says to plan for both.

  1. Spec against the model you will ship on. Capability forecasts from the best informed people available have run years slow, so keep the harness thin enough that a stronger model drops in without a rewrite.
  2. Plan the edges on the slow clock. Human review, sign off, integrations with physical systems and inexperienced operators set the pace of delivered value, and a faster model does not move them.
  3. Measure what comes out the far side of the operator. The trial shows that delivered capability can sit far below what a benchmark implies, and that the people who know the domain best will guess too high.

The next forecaster runs on the fast clock

None of this is settled, and FRI says so more clearly than most would. Its method favours finding underestimates: a forecast that is too low is exposed the moment reality crosses it, while one that is too high can only be confirmed after its resolution date passes. Most questions have not resolved. It also works from medians, which hide the subsets of forecasters who did much better. The overestimates may be waiting in 2030, and the questions that matter most, on growth, labour, inequality and catastrophic risk, are the ones that resolve last.

Where questions had not resolved, FRI judged the humans against a model's projections, and it flags openly that the model had months of hindsight. ForecastBench already shows some models forecasting at superforecaster level on some question types, and FRI's next step is to add continuously updated LLM forecasts to the panel itself. The logic is sound: if people cannot keep pace with the curve, use something that reads it in real time.

But the new forecaster lives on the fast side of the seam, and the questions it will be asked about sit on the slow one. The experts who got benchmarks wrong by years got autonomous vehicles wrong in the other direction. Nobody yet knows which of those mistakes the machine will make.

Unlock the Future of Business with AI

Dive into our immersive workshops and equip your team with the tools and knowledge to lead in the AI era.

Scroll to top