Back

THE ARCHIVE

Every fact, every receipt

Every number we publish, with the body that produced it and the date it was true. The tier on each card says how strong the receipt is — straight from the source, reported second-hand, or a named analyst's estimate. Nothing here is unsourced. How we source these.

  • Primarystraight from the source
  • Reportedvia a named outlet
  • Estimatea named analyst's projection
15 facts
ModelsReported

Frontier AI models are catching up fast to benchmarks meant to resist them. A year ago they scored under 10% on Humanity's Last Exam, a 2,700-question test designed to be hard for AI and favorable to human experts. Now accuracy has reached 38.3% on Humanity's Last Exam, up from under 10% a year earlier

SourceStanford AI Index 2026, citing CAIS·2026 AI Index Report
Open this fact
ModelsPrimary

Epoch AI's Capabilities Index scores frontier models on one aggregate scale, and fits a break in the trend at April 2024, when reasoning models and reinforcement learning took over frontier training. Before the break the frontier gained 8.3 index points a year; after it, 15.5 points a year, 1.85 times as fast

SourceEpoch AI·December 2025
Open this fact
ModelsPrimary

Epoch AI tested three commercial AI detectors — Pangram, GPTZero, Originality.ai — pinned to their June 2026 versions. Against plain machine-generated text they catch nearly everything, but shown five samples of a specific author's writing and told to imitate the style, AI detectors missed 13% of style-imitated AI passages, versus under 1% of plain AI text

SourceEpoch AI·Jul 2026
Open this fact
ModelsPrimary

Epoch AI fits a trend through the models whose context windows ranked among the ten longest on their release date, and finds that since mid-2023 the frontier context window has been getting about 30× longer every year

SourceEpoch AI·since mid-2023
Open this fact
ModelsPrimary

Frontier models have started finding software flaws at scale. Counting the high- and critical-severity CVEs disclosed by 21 major vendors — Microsoft, Google, Apple and Cisco among them — Epoch AI puts July 2026 at about 2,500 serious CVEs in July 2026, five times the monthly record set before April 2026

SourceEpoch AI·July 2026
Open this fact
ModelsPrimary

Epoch AI tracked 333 compute estimates for notable AI models released between 2010 and May 2024, and found that training compute for the frontier — the top 10 models by compute at release — has been growing each year to reach 5.3x the prior year's level

SourceEpoch AI·2010–May 2024 data
Open this fact
ModelsPrimary

Epoch AI audited its FrontierMath benchmark after OpenAI flagged suspiciously high scores, and its 12 June 2026 v2 release corrected errors in 42% of the original problems, leaving a verified set of 338 problems across four tiers 42% of FrontierMath's original problems contained errors

SourceEpoch AI·June 2026 (v2 release)
Open this fact
ModelsEstimate

AI capability is getting cheaper to reach, not just bigger. Epoch AI tracked what it cost in tokens to hit the same score — about 27% accuracy on its FrontierMath benchmark — first with o4-mini in April 2025, then with GPT-5.2 in December 2025, and found AI capability got roughly 3x cheaper to reach in eight months, part of a ~5-10x per year cost-decline trend

SourceEpoch AI·Feb 2026
Open this fact
ModelsEstimate

Across 42 notable language models, Epoch AI estimates frontier training costs have risen from roughly $2 million for GPT-3 to up to nearly $390 million per run

SourceEpoch AI·2024
Open this fact
ModelsPrimary

Epoch AI's Open Problems benchmark collects research mathematics problems that have resisted serious attempts by professional mathematicians. Expanded to 50 problems on 31 July 2026, the count two weeks later — with the newest AI solution still marked provisional — stood at 4 solved by AI, 1 by a human, 45 still open

SourceEpoch AI FrontierMath·17 August 2026
Open this fact
ModelsPrimary

Epoch AI's Capabilities Index scores language models on one aggregate scale, and nothing since has held first place as long as GPT-4 did after its March 2023 release — roughly a year. The second-longest lead, by OpenAI's o1, lasted a little over three months, less than a third of GPT-4's run at the top

Open this fact
ModelsPrimary

Across January to May 2026, Epoch AI's capabilities index puts the best open-weight models behind the frontier closed models by an average of four months

SourceEpoch AI·Jan-May 2026
Open this fact
ModelsPrimary

Scale is not the only driver of AI progress. Epoch AI measures how much training compute it takes to reach a fixed level of language-model performance, and finds that requirement falling fast: pre-training compute efficiency doubles every 7.6 months

SourceEpoch AI·Feb 2026
Open this fact