Back
Models6/day by emailPrimary

Stanford's 2026 AI Index tracks how AI models perform on GPQA Diamond, a set of PhD-level chemistry and physics questions designed to resist Google searches. Mean model accuracy on the benchmark reached 93% in 2025, compared with an expert human validator baseline of 81.2% -- AI models now beat PhD-level experts by 12 points on GPQA Diamond

SourceStanford HAI / Epoch AI·2026 AI Index

Verified 2026-10-01

THE INTELLIGENCE CLUB

Facts like this, six days a week

The Daily Receipts on weekdays, The Weekly Brief on Saturday — six sourced facts a day by email, two more than the free site shows. Every one traced to the filing it came from. Free.

Join now and you start with 100 Receipts on The Receipts Ledger — your first step toward a free year of Insider or Pro. How it works →

Read this morning's edition →