Back
Models6/day by emailReported
Stanford's 2026 AI Index cites a study that pitted Microsoft's AI Diagnostic Orchestrator, paired with OpenAI's o3 reasoning model, against 21 practicing physicians with five to twenty years of experience working unaided on diagnostically challenging New England Journal of Medicine cases. The AI system's accuracy reached 85.5%, versus about 20% for the unaided physicians
Verified 2026-09-28
THE INTELLIGENCE CLUB
Facts like this, six days a week
The Daily Receipts on weekdays, The Weekly Brief on Saturday — six sourced facts a day by email, two more than the free site shows. Every one traced to the filing it came from. Free.
Free forever, but only while founding spots last. Get in now, before this is just a newsletter everyone's on.
Read this morning's edition →Microsoft’s exposure cardWhat share of the business the AI build-out actually isThe screener · 72 companiesEvery layer of the AI chain, priced dailyHow this was verifiedEvery figure links to the filing it came from
More on Models
- AI models can now win gold at the International Mathematical Olympiad, yet Stanford's 2026 AI Index found they still fumble a task most children master early. Across ClockBench's 180 clock designs and 720 questions, the top model, GPT-5.4 High, read analog clocks correctly only 50.6% of the time, versus 90.1% for humans
- Frontier AI models are catching up fast to benchmarks meant to resist them. A year ago they scored under 10% on Humanity's Last Exam, a 2,700-question test designed to be hard for AI and favorable to human experts. Now accuracy has reached 38.3% on Humanity's Last Exam, up from under 10% a year earlier
- AI agents attempting OSWorld's real-world computer tasks across Ubuntu, Windows and macOS — file operations, multi-app workflows — historically topped out at just 1% to 12% success. Stanford's 2026 AI Index reports the best model, Claude Opus 4.5, now reaches 66.3% accuracy, within 6 percentage points of the 72.35% human baseline 66.3% accuracy on OSWorld's real-computer-task benchmark