Back
ModelsPrimary

SWE-bench Verified hands an AI model a real GitHub issue and a codebase and checks whether the patch it writes actually fixes the bug. Stanford's 2026 AI Index reports the leading model, Claude 4.5 Opus in high-reasoning mode, had solved about 76.8% of the benchmark's issues as of February 2026

SourceStanford AI Index·Feb 2026

Verified 2026-09-04

THE INTELLIGENCE CLUB

Facts like this, six days a week

The Daily Receipts on weekdays, The Weekly Brief on Saturday — every fact traced to the filing it came from. Free.

Free forever. Six emails a week — The Daily Receipts on weekdays, The Weekly Brief on Saturday — drop either in one click.

Read this morning's edition →