Back
ModelsReported

Terminal-Bench 2.0 tests whether AI agents can chain together the kind of multi-step, unsupervised terminal work a developer does in a day — compiling code, training models, setting up servers. Stanford's 2026 AI Index reports agent accuracy on it nearly quadrupled in a year, reaching 77.3% by early 2026, up from 20% in February 2025

Verified 2026-09-02

THE INTELLIGENCE CLUB

Facts like this, six days a week

The Daily Receipts on weekdays, The Weekly Brief on Saturday — every fact traced to the filing it came from. Free.

Free forever. Six emails a week — The Daily Receipts on weekdays, The Weekly Brief on Saturday — drop either in one click.

Read this morning's edition →