Back
ModelsPrimary
Epoch AI's Open Problems benchmark collects research mathematics problems that have resisted serious attempts by professional mathematicians. Expanded to 50 problems on 31 July 2026, the count two weeks later — with the newest AI solution still marked provisional — stood at 4 solved by AI, 1 by a human, 45 still open
Verified 2026-08-17
THE INTELLIGENCE CLUB
Facts like this, six days a week
The Daily Receipts on weekdays, The Weekly Brief on Saturday — every fact traced to the filing it came from. Free.
Free forever. Six emails a week — The Daily Receipts on weekdays, The Weekly Brief on Saturday — drop either in one click.
Read this morning's edition →The screener · 72 companiesEvery layer of the AI chain, priced dailyHow this was verifiedEvery figure links to the filing it came from
More on Models
- Frontier AI models are catching up fast to benchmarks meant to resist them. A year ago they scored under 10% on Humanity's Last Exam, a 2,700-question test designed to be hard for AI and favorable to human experts. Now accuracy has reached 38.3% on Humanity's Last Exam, up from under 10% a year earlier
- Epoch AI's Capabilities Index scores frontier models on one aggregate scale, and fits a break in the trend at April 2024, when reasoning models and reinforcement learning took over frontier training. Before the break the frontier gained 8.3 index points a year; after it, 15.5 points a year, 1.85 times as fast
- Epoch AI tested three commercial AI detectors — Pangram, GPTZero, Originality.ai — pinned to their June 2026 versions. Against plain machine-generated text they catch nearly everything, but shown five samples of a specific author's writing and told to imitate the style, AI detectors missed 13% of style-imitated AI passages, versus under 1% of plain AI text