Home Report A2H Protocol Books Newsletter Articles About
The Intelligence Advantage · Report Update
The H1 2026 Model Shift

The H1 2026 Model Shift

How frontier models and model intelligence changed between this report's original March 2026 edition (data as of February 28, 2026) and the updated July 2026 edition. A five-month view across the material of Chapters 1, 3, and 4.

February 2026 → July 2026
Frontier SWE-bench
82.1%~95.5%
Best verified coding score. Mythos-class now near-solves the benchmark.
Open-source GPQA
75%91.2%
Best open-weight score on GPQA Diamond — now within ~3 pts of the frontier.
Frontier → open gap
~6 mo~4 mo
Epoch AI capability lag. Coding gap now roughly 1–2 months.
Cheapest capable
$0.27$0.14
Per M input tokens. DeepSeek V3.2 → V4 — the 1,000x curve kept going.

The original edition of this report captured a market in mid-motion: a single February wave of releases — GPT-5.2, Claude Opus 4.6, Gemini 3.1 Pro, Grok 4.1, GLM-5, Kimi K2.5, Qwen 3, and DeepSeek V3.2 — that the report described as "seven major releases in one month." Five months later, essentially that entire roster has turned over. This page holds the two editions side by side and asks a narrow question: what actually moved in the market between February 28 and mid-July 2026? The numbers below are pulled directly from both editions of Chapters 1, 3, and 4; unverifiable July figures are marked est.

1 The Release Wave

One full generation of frontier and open-weight models turned over in roughly five months. The left column is what the original edition captured; the right is what arrived after it went out.

Captured in the original edition

The February 2026 wave — "seven major releases in one month"
Feb 2026
GPT-5.2 (OpenAI) — AIME 100%, the closed reasoning flagship
Feb 2026
Claude Opus 4.6 (Anthropic) — #1 Arena; Sonnet 5 at 82.1% SWE-bench
Feb 2026
Gemini 3.1 Pro (Google) — GPQA 94.3%, ARC-AGI-2 77.1%
Feb 2026
Grok 4.1 (xAI)
Feb 2026
GLM-5 (Zhipu, MIT) — led HLE among open weights
Feb 2026
Kimi K2.5 (Moonshot) — 200–300 autonomous tool calls
Feb 2026
Qwen 3 (Alibaba) & DeepSeek V3.2 — the cheap-capable open tier

Arrived since (March–July 2026)

What the updated edition adds
Apr 24
DeepSeek V4 (Pro / Flash) — new cheapest-capable tier at $0.14/M
Jun 30
Claude Sonnet 5 (Anthropic); Opus 4.8 as the flagship tier
Jun/Jul
GPT-5.6 family (Sol / Terra / Luna) — ChatGPT default from ~Jul 9
Jul
Grok 4.5 (xAI) — supersedes Grok 4.1
Jul 16
Kimi K3 (Moonshot) — ~2.8T parameters
H1 2026
GLM-5.2 (#1 open-weight, MIT), Qwen 3.6, Mistral Large 3
Mythos-class models (Fable 5 / Mythos 5) now top the hardest coding benchmarks
One generation of turnover in five months — the whole February roster has a named successor.
2 Benchmark Movement

Chapter 3 and 4 material. Two benchmarks moved hard — frontier coding and open-weight science — while the rest were already saturated or remain stubbornly unsolved.

February vs. July 2026 — best score by benchmark

Grouped bars: the darker bar is July. Frontier SWE-bench and open-weight GPQA jumped; frontier GPQA, MMLU, AIME, and HLE barely moved.
BenchmarkMarch edition (Feb 28)July editionWhat moved
SWE-bench Verified 82.1% — Claude Sonnet 5 (frontier best) ~95.5% Mythos 5 / 95.0% Fable 5; Opus 4.8 ~88.6%, Sonnet 5 ~85.2% +13 pts Real-world coding went from "hard" to nearly solved at the frontier.
GPQA Diamond (frontier) 94.3% — Gemini 3.1 Pro 94.3% — Gemini 3.1 Pro (still leads) flat Frontier science already near a natural ceiling.
GPQA Diamond (open-weight) ~75% — best open model 91.2% — GLM-5.2 (MIT) +16 pts The single biggest convergence move — now ~3 pts off the frontier.
HLE (Humanity's Last Exam) ~50.4% best (GLM-5 led open); all near 50% Every model still below 51% flat The hardest differentiator; essentially unmoved in five months.
MMLU Saturated ~92%; open-closed gap ~0.3 pt Saturated ~92%+; no longer a differentiator flat Retired as a discriminator on both dates.
AIME 100% — GPT-5.2 100% — GPT-5.6 family flat Competition math saturated; only the model name changed.
3 The Pricing Shift

Chapter 1 and 4 material. The full published range held at roughly 1,000x ($0.02–$21/M) across both editions — but the floor kept dropping, and the tiers reshuffled as new families landed. H1 continued the intelligence-per-dollar curve rather than bending it.

TierMarch edition (input $/M)July edition (input $/M)What it does to intelligence-per-dollar
Cheapest capable DeepSeek V3.2 — $0.27 DeepSeek V4 — $0.14 Floor roughly halved. V4 is ~36x cheaper than the Opus 4.8 flagship for strong general + math work.
#1 open-weight GLM-5 — $0.11 corrected GLM-5.2 — $1.40 (MIT, 1M ctx) The Feb "$0.11" was unsupported; the July edition prices the real #1 open tier at $1.40 (cached ~$0.26).
Mid-tier Gemini 3 Flash $0.50; GPT-5.2 std $2.50 Gemini Flash-Lite $0.10; GPT-5.6 Luna $1.00 A genuinely capable mid-tier landed well under $1/M, compressing the middle of the curve.
Frontier flagship Claude Opus 4.5/4.6 — $5.00 Opus 4.8 / GPT-5.6 Sol — $5.00 Flagship list price held flat at $5/M — the frontier price anchor did not move.
Reasoning premium GPT-5.2 reasoning $15 (up to ~$21) Frontier Pro reasoning ~$21 est. The margin has moved from raw intelligence to compute-time reasoning — a 10–15x premium on both dates.

The through-line for Chapter 1's thesis: cheaper intelligence did not slow consumption. Published token prices dropped ~1,000x over the arc while total AI inference spending surged 320% in 2025 — the Jevons flywheel the report describes. H1 2026's contribution was to push the cheap-capable floor from $0.27 to $0.14 while leaving the frontier anchor at $5, widening the intelligence-per-dollar spread rather than compressing it.

4 The Convergence Story

Chapter 3's central claim — that the open/closed gap is collapsing — got measurably stronger between the two editions, but not uniformly. Where the frontier still leads, it leads on the hardest problems.

~4 months

The gap narrowed

Epoch AI's frontier-vs-open capability lag went from "under six months" in the original edition to roughly four months (range 3–7) by July. On coding specifically, the gap is now about 1–2 months.

~3 points

Open weights nearly caught the frontier

GLM-5.2 posts 91.2% on GPQA Diamond against the frontier's 94.3% — open-weight science is now within about three points, up from a ~19-point gap earlier in the arc.

<51%

What stayed closed-led

HLE remains unsolved by every model, and the very hardest coding still belongs to closed Mythos-class systems. The convergence is real everywhere except the frontier's frontier.

About this comparison

This page compares market reality between two dates — February 28, 2026 and mid-July 2026 — not editorial corrections. The July edition also fixed several errors carried in the February edition: the GPQA human-expert baseline (~65% per Rein et al. 2023, not the ~90% figure that was actually a model score), a Grok version label, and GLM-5's pricing (the "$0.11/M" figure was unsupported and has been re-anchored to the real #1 open tier). Those are editorial fixes, and they are deliberately kept out of the "what moved" columns above so the comparison reflects the market, not the copyediting.

Where a July value could not be independently verified — principally the frontier reasoning price tier — it is marked est. and treated as an editorial estimate. Frontier SWE-bench is shown as ~95.5% (Mythos-class); the published Chapter 4 comparison lists Opus 4.8 at 88.6% and Sonnet 5 at 85.2%, so the headline reflects the best available model, not the median flagship.

The Intelligence Advantage — a report by Jose Antonio Martinez Aguilar. Published July 2026. Figures drawn from the report's February 2026 and July 2026 editions; unverifiable July cells marked (est.).