Video generation models can now produce visually convincing footage, but Stanford's 2026 AI Index finds they still struggle to get the physics and consistency right. On VBench-2.0, which scores videos for human-aligned faithfulness to reality rather than visual polish alone, no model evaluated in early 2026 topped a total score of Veo 3 leads the VBench-2.0 video-realism benchmark at 66.7% — no AI video model tested tops 67%
Verified 2026-09-06
THE INTELLIGENCE CLUB
Facts like this, six days a week
The Daily Receipts on weekdays, The Weekly Brief on Saturday — every fact traced to the filing it came from. Free.
Free forever. Six emails a week — The Daily Receipts on weekdays, The Weekly Brief on Saturday — drop either in one click.
Read this morning's edition →More on Models
- Frontier AI models are catching up fast to benchmarks meant to resist them. A year ago they scored under 10% on Humanity's Last Exam, a 2,700-question test designed to be hard for AI and favorable to human experts. Now accuracy has reached 38.3% on Humanity's Last Exam, up from under 10% a year earlier
- AI agents attempting OSWorld's real-world computer tasks across Ubuntu, Windows and macOS — file operations, multi-app workflows — historically topped out at just 1% to 12% success. Stanford's 2026 AI Index reports the best model, Claude Opus 4.5, now reaches 66.3% accuracy, within 6 percentage points of the 72.35% human baseline 66.3% accuracy on OSWorld's real-computer-task benchmark
- SWE-bench Verified hands an AI model a real GitHub issue and a codebase and checks whether the patch it writes actually fixes the bug. Stanford's 2026 AI Index reports the leading model, Claude 4.5 Opus in high-reasoning mode, had solved about 76.8% of the benchmark's issues as of February 2026