Back
ModelsReported

AI agents attempting OSWorld's real-world computer tasks across Ubuntu, Windows and macOS — file operations, multi-app workflows — historically topped out at just 1% to 12% success. Stanford's 2026 AI Index reports the best model, Claude Opus 4.5, now reaches 66.3% accuracy, within 6 percentage points of the 72.35% human baseline 66.3% accuracy on OSWorld's real-computer-task benchmark

Verified 2026-08-31

THE INTELLIGENCE CLUB

Facts like this, six days a week

The Daily Receipts on weekdays, The Weekly Brief on Saturday — every fact traced to the filing it came from. Free.

Free forever. Six emails a week — The Daily Receipts on weekdays, The Weekly Brief on Saturday — drop either in one click.

Read this morning's edition →