
Frontier AI models are catching up fast to benchmarks meant to resist them. A year ago they scored under 10% on Humanity's Last Exam, a 2,700-question test designed to be hard for AI and favorable to human experts. Now accuracy has reached 38.3% on Humanity's Last Exam, up from under 10% a year earlier
Open this fact →