Back
Models6/day by emailPrimary
Epoch AI compared how OpenAI's and Anthropic's flagship models handle very long prompts, and found GPT-5.6 Terra and Sol slow down quadratically as context grows while Claude Sonnet 5 and Opus 5 stay almost linear -- the marginal latency to process the next 10,000 tokens for GPT-5.6 Terra rose from 0.071 seconds at a 100,000-token prompt to 1.66 seconds at a 10-million-token prompt
Verified 2026-09-11
THE INTELLIGENCE CLUB
Facts like this, six days a week
The Daily Receipts on weekdays, The Weekly Brief on Saturday — six sourced facts a day by email, two more than the free site shows. Every one traced to the filing it came from. Free.
Free forever. Six emails a week — The Daily Receipts on weekdays, The Weekly Brief on Saturday — drop either in one click.
Read this morning's edition →The screener · 72 companiesEvery layer of the AI chain, priced dailyHow this was verifiedEvery figure links to the filing it came from
More on Models
- AI models can now win gold at the International Mathematical Olympiad, yet Stanford's 2026 AI Index found they still fumble a task most children master early. Across ClockBench's 180 clock designs and 720 questions, the top model, GPT-5.4 High, read analog clocks correctly only 50.6% of the time, versus 90.1% for humans
- Frontier AI models are catching up fast to benchmarks meant to resist them. A year ago they scored under 10% on Humanity's Last Exam, a 2,700-question test designed to be hard for AI and favorable to human experts. Now accuracy has reached 38.3% on Humanity's Last Exam, up from under 10% a year earlier
- AI agents attempting OSWorld's real-world computer tasks across Ubuntu, Windows and macOS — file operations, multi-app workflows — historically topped out at just 1% to 12% success. Stanford's 2026 AI Index reports the best model, Claude Opus 4.5, now reaches 66.3% accuracy, within 6 percentage points of the 72.35% human baseline 66.3% accuracy on OSWorld's real-computer-task benchmark