How frontier models and model intelligence changed between this report's original March 2026 edition (data as of February 28, 2026) and the updated July 2026 edition. A five-month view across the material of Chapters 1, 3, and 4.
The original edition of this report captured a market in mid-motion: a single February wave of releases — GPT-5.2, Claude Opus 4.6, Gemini 3.1 Pro, Grok 4.1, GLM-5, Kimi K2.5, Qwen 3, and DeepSeek V3.2 — that the report described as "seven major releases in one month." Five months later, essentially that entire roster has turned over. This page holds the two editions side by side and asks a narrow question: what actually moved in the market between February 28 and mid-July 2026? The numbers below are pulled directly from both editions of Chapters 1, 3, and 4; unverifiable July figures are marked est.
One full generation of frontier and open-weight models turned over in roughly five months. The left column is what the original edition captured; the right is what arrived after it went out.
Chapter 3 and 4 material. Two benchmarks moved hard — frontier coding and open-weight science — while the rest were already saturated or remain stubbornly unsolved.
| Benchmark | March edition (Feb 28) | July edition | What moved |
|---|---|---|---|
| SWE-bench Verified | 82.1% — Claude Sonnet 5 (frontier best) | ~95.5% Mythos 5 / 95.0% Fable 5; Opus 4.8 ~88.6%, Sonnet 5 ~85.2% | +13 pts Real-world coding went from "hard" to nearly solved at the frontier. |
| GPQA Diamond (frontier) | 94.3% — Gemini 3.1 Pro | 94.3% — Gemini 3.1 Pro (still leads) | flat Frontier science already near a natural ceiling. |
| GPQA Diamond (open-weight) | ~75% — best open model | 91.2% — GLM-5.2 (MIT) | +16 pts The single biggest convergence move — now ~3 pts off the frontier. |
| HLE (Humanity's Last Exam) | ~50.4% best (GLM-5 led open); all near 50% | Every model still below 51% | flat The hardest differentiator; essentially unmoved in five months. |
| MMLU | Saturated ~92%; open-closed gap ~0.3 pt | Saturated ~92%+; no longer a differentiator | flat Retired as a discriminator on both dates. |
| AIME | 100% — GPT-5.2 | 100% — GPT-5.6 family | flat Competition math saturated; only the model name changed. |
Chapter 1 and 4 material. The full published range held at roughly 1,000x ($0.02–$21/M) across both editions — but the floor kept dropping, and the tiers reshuffled as new families landed. H1 continued the intelligence-per-dollar curve rather than bending it.
| Tier | March edition (input $/M) | July edition (input $/M) | What it does to intelligence-per-dollar |
|---|---|---|---|
| Cheapest capable | DeepSeek V3.2 — $0.27 | DeepSeek V4 — $0.14 | Floor roughly halved. V4 is ~36x cheaper than the Opus 4.8 flagship for strong general + math work. |
| #1 open-weight | GLM-5 — $0.11 corrected | GLM-5.2 — $1.40 (MIT, 1M ctx) | The Feb "$0.11" was unsupported; the July edition prices the real #1 open tier at $1.40 (cached ~$0.26). |
| Mid-tier | Gemini 3 Flash $0.50; GPT-5.2 std $2.50 | Gemini Flash-Lite $0.10; GPT-5.6 Luna $1.00 | A genuinely capable mid-tier landed well under $1/M, compressing the middle of the curve. |
| Frontier flagship | Claude Opus 4.5/4.6 — $5.00 | Opus 4.8 / GPT-5.6 Sol — $5.00 | Flagship list price held flat at $5/M — the frontier price anchor did not move. |
| Reasoning premium | GPT-5.2 reasoning $15 (up to ~$21) | Frontier Pro reasoning ~$21 est. | The margin has moved from raw intelligence to compute-time reasoning — a 10–15x premium on both dates. |
The through-line for Chapter 1's thesis: cheaper intelligence did not slow consumption. Published token prices dropped ~1,000x over the arc while total AI inference spending surged 320% in 2025 — the Jevons flywheel the report describes. H1 2026's contribution was to push the cheap-capable floor from $0.27 to $0.14 while leaving the frontier anchor at $5, widening the intelligence-per-dollar spread rather than compressing it.
Chapter 3's central claim — that the open/closed gap is collapsing — got measurably stronger between the two editions, but not uniformly. Where the frontier still leads, it leads on the hardest problems.
Epoch AI's frontier-vs-open capability lag went from "under six months" in the original edition to roughly four months (range 3–7) by July. On coding specifically, the gap is now about 1–2 months.
GLM-5.2 posts 91.2% on GPQA Diamond against the frontier's 94.3% — open-weight science is now within about three points, up from a ~19-point gap earlier in the arc.
HLE remains unsolved by every model, and the very hardest coding still belongs to closed Mythos-class systems. The convergence is real everywhere except the frontier's frontier.
This page compares market reality between two dates — February 28, 2026 and mid-July 2026 — not editorial corrections. The July edition also fixed several errors carried in the February edition: the GPQA human-expert baseline (~65% per Rein et al. 2023, not the ~90% figure that was actually a model score), a Grok version label, and GLM-5's pricing (the "$0.11/M" figure was unsupported and has been re-anchored to the real #1 open tier). Those are editorial fixes, and they are deliberately kept out of the "what moved" columns above so the comparison reflects the market, not the copyediting.
Where a July value could not be independently verified — principally the frontier reasoning price tier — it is marked est. and treated as an editorial estimate. Frontier SWE-bench is shown as ~95.5% (Mythos-class); the published Chapter 4 comparison lists Opus 4.8 at 88.6% and Sonnet 5 at 85.2%, so the headline reflects the best available model, not the median flagship.