Back
ModelsPrimary

Epoch AI audited its FrontierMath benchmark after OpenAI flagged suspiciously high scores, and its 12 June 2026 v2 release corrected errors in 42% of the original problems, leaving a verified set of 338 problems across four tiers 42% of FrontierMath's original problems contained errors

SourceEpoch AI·June 2026 (v2 release)

Verified 2026-08-24

THE INTELLIGENCE CLUB

Facts like this, six days a week

The Daily Receipts on weekdays, The Weekly Brief on Saturday — every fact traced to the filing it came from. Free.

Free forever. Six emails a week — The Daily Receipts on weekdays, The Weekly Brief on Saturday — drop either in one click.

Read this morning's edition →