Frontier AI Models — Comprehensive Comparison

Commercial & Open-Source Large Language Models | Benchmarks, Capabilities, Pricing & Distance Analysis
July 2026 · orig. February 2026
Gemini 3.1 Pro
Claude Opus 4.8
Claude Sonnet 5
GPT-5.6
Grok 4.5
GLM-5.2
Kimi K3
DeepSeek V4
Qwen 3.6
OLMo 3.1
Llama 4
Part I — The Intelligence Layer
Chapter 4: Frontier Model Comparison

This companion analysis breaks down the current model landscape by tier, mapping the specific strengths, weaknesses, and cost profiles of every frontier model available in July 2026. Where the primary convergence analysis tracks how the gap has closed, this page answers the practical question: which models belong in which tier, and why. Composite scores and any unpublished benchmark cells are editorial estimates.

The tier framework below classifies models into four levels — S, A, B, and C — based on aggregate benchmark performance, real-world capability, and competitive positioning. The critical finding: S-tier is no longer the exclusive province of closed models. Open-weight models now occupy A-tier convincingly, and on select benchmarks, individual open models outperform every closed alternative. For the detailed convergence trajectory and economic analysis, see Chapter 3: The Great Convergence.

T Model Tier Ranking — Overall Frontier Capability
S-TIER — Frontier Leaders
Gemini 3.1 Pro
Google DeepMind · current
Closed
GPQA: 94.3% · #1 reasoning · ARC-AGI-2 leader (est.)
Claude Opus 4.8
Anthropic · current
Closed NEW
SWE-bench: 88.6% · Arena #1 · Best all-rounder
GPT-5.6
OpenAI · Jun/Jul 2026
Closed NEW
Sol / Terra / Luna tiers · strong math · low hallucination
Claude Sonnet 5
Anthropic · Jun 30 2026
Closed NEW
SWE-bench: 85.2% · $2/MTok value · Mythos-class tops coding (~95.5%)
A-TIER — Elite Contenders
Grok 4.5
xAI · Jul 2026
Closed NEW
Strong math · AIME: 100% · Arena top-5 · $2/MTok
GLM-5.2
Zhipu AI · current
Open MIT NEW
GPQA: 91.2% · #1 open-source · 1M ctx · $1.40/MTok
Kimi K3
Moonshot AI · Jul 16 2026
Open NEW
~2.8T params · best agentic open model · HLE <51%
DeepSeek V4
DeepSeek · Apr 2026
Open NEW
$0.14/MTok · cheapest capable · strong math
B-TIER — Strong Performers
Qwen 3.6
Alibaba · 2026
Open Apache 2.0
700M+ downloads · Most used open model · 201 languages
Mistral Large 3
Mistral AI · 2026
Open weight
Large 3 / Medium 3.5 / Small 4 · EU frontier lab
Llama 4
Meta · 2025
Open
10M context · 109B MoE · GPT-4o era performance
OLMo 3.1
AI2 (Allen Institute) · 2025
Fully Open
Full transparency · 32B · Best for research

Reading the Tier Map

S-tier models — Gemini 3.1 Pro, Claude Opus 4.8, and GPT-5.6 — represent the current frontier ceiling. No single model dominates every benchmark; instead, each claims specific territory. Gemini leads reasoning (GPQA: 94.3%), Claude leads coding and human preference (SWE-bench: 88.6% for Opus 4.8, ~95.5% for Anthropic’s Mythos-class coding specialists, #1 Arena Elo), and GPT-5.6 anchors mathematics and factuality. The A-tier is where convergence becomes tangible: GLM-5.2, Kimi K3, and Grok 4.5 deliver S-tier-adjacent performance at dramatically lower cost — or in the case of open models, at near-zero marginal cost when self-hosted.

B Benchmark Comparison — Key Scores
ARC-AGI-2 Novel problem solving & fluid reasoning · scores are editorial estimates (July leaders unverified)
Gemini 3.1 Pro
~77%
Claude Opus 4.8
~69%
Claude Sonnet 5
~58%
GPT-5.6 Pro
~54%
Gemini 3.1 Deep Think
~45%
GPQA Diamond Graduate-level science · PhD experts ~65% (Rein et al. 2023)
Gemini 3.1 Pro
94.3%
GLM-5.2 (Open)
91.2%
Claude Opus 4.8
~91% (est.)
Grok 4.5
~89% (est.)
Claude Sonnet 5
~85% (est.)
Llama 4
69.8%
SWE-bench Verified Real-world software engineering · open-weight and non-Anthropic July scores are editorial estimates
Claude Mythos 5
95.5%
Claude Fable 5
95.0%
Claude Opus 4.8
88.6%
GLM-5.2 (Open)
~88% (est.)
Kimi K3 (Open)
~87% (est.)
Claude Sonnet 5
85.2%
GPT-5.6
~85% (est.)
Gemini 3.1 Pro
~82% (est.)
AIME 2025 Competition mathematics · period scores; no newer AIME result verified for July
GPT-5.6
100%
Grok 4.5
100%
Kimi K3
96.1%
DeepSeek V4
96.0%
Gemini 3.1 Pro
~95%
Qwen 3.6 Think
92%
Humanity's Last Exam (w/ tools) Hardest multi-domain exam · every model still below 51%; leaders contested (est.)
GLM-5.2 (Open)
~50%
Kimi K3 (Open)
~50%
GPT-5.6
~46%
Claude Opus 4.8
~43%
Gemini 3.1 Deep Think
~41%

Beyond Single Scores

Benchmarks capture isolated capabilities, but enterprise deployment demands multi-dimensional strength. The capability heatmap below reveals what aggregate scores obscure: models that appear similar on headline numbers often diverge sharply in domain-specific performance. This is why model routing — matching the right model to the right task — matters more than selecting a single “best” model.

M Capability Heatmap — Relative Strength by Domain
Scale: 5 = Best in class   4 = Excellent   3 = Strong   2 = Average   1 = Below avg   = No data
Capability Gemini
3.1 Pro
Claude
Opus 4.8
Claude
Sonnet 5
GPT
5.6
Grok
4.5
GLM-5.2 Kimi
K3
DS
V4
Qwen
3.6
Llama
4
OLMo
3.1
REASONING & KNOWLEDGE
Logical Reasoning 5 4 4 4 4 3 3 3 3 2 2
Scientific Knowledge 5 5 3 4 4 3 3 3 3 2 2
Mathematics 5 4 4 5 5 3 5 5 4 2 3
CODING & ENGINEERING
Code Generation 3 5 5 5 3 4 4 3 3 2 2
Bug Fixing (SWE) 3 5 5 5 3 4 4 3 3 2 2
AGENTIC & TOOL USE
Tool Calling 4 5 5 5 4 4 5 3 3 2 2
Multi-step Agents 4 5 5 5 4 4 5 3 3 2 1
LANGUAGE & CREATIVITY
Creative Writing 4 5 4 4 5 3 3 2 3 3 2
Multilingual 4 4 4 4 3 4 3 3 5 4 2
MULTIMODAL
Vision/Image 5 4 3 4 3 3 5 2 4 3 1
RELIABILITY
Low Hallucination 4 4 4 5 4 2 2 3 3 3 3
Long Context 5 5 5 4 4 3 3 3 3 5 2
Instruction Following 5 5 5 5 4 4 4 3 4 3 3
Multi-Dimensional Capability Profiles (0-100 composite)
Gemini 3.1 Pro
Reasoning
97
Math
95
Coding
76
Agentic
82
Vision
92
Reliability
88
Claude Opus 4.8
Reasoning
91
Math
88
Coding
96
Agentic
95
Vision
82
Reliability
92
GPT-5.6
Reasoning
88
Math
98
Coding
92
Agentic
93
Vision
84
Reliability
96
Grok 4.5
Reasoning
86
Math
97
Coding
72
Agentic
78
Vision
68
Reliability
85
GLM-5.2 (Open)
Reasoning
78
Math
72
Coding
85
Agentic
82
Vision
68
Reliability
65
Kimi K3 (Open)
Reasoning
75
Math
94
Coding
80
Agentic
95
Vision
90
Reliability
55
DeepSeek V4 (Open)
Reasoning
78
Math
98
Coding
70
Agentic
60
Vision
45
Reliability
68
Qwen 3.6 (Open)
Reasoning
72
Math
88
Coding
75
Agentic
68
Vision
80
Reliability
70
R Composite Power Ranking — Weighted Average Across All Benchmarks
Weighted composite: Reasoning (20%) + Math (15%) + Coding (20%) + Agentic (15%) + Vision (10%) + Reliability (10%) + Human Preference (10%)
#1
Claude Opus 4.8
Anthropic93
#2
Gemini 3.1 Pro
Google91
#3
GPT-5.6
OpenAI91
#4
Claude Sonnet 5
Anthropic89
#5
Grok 4.5
xAI84
#6
Kimi K3
Moonshot · Open82
#7
GLM-5.2
Zhipu · Open MIT79
#8
DeepSeek V4
DeepSeek · Open75
#9
Qwen 3.6
Alibaba · Open Apache74
#10
Llama 4
Meta · Open64
#11
OLMo 3.1
AI2 · Fully Open52

The Cost-Capability Equation

The pricing table below makes the strategic calculus explicit. DeepSeek V4 offers near-frontier capability at $0.14/MTok input — roughly 36x cheaper than Claude Opus 4.8 ($5) and about 150x cheaper than a $21 frontier reasoning tier. GLM-5.2, the #1 open-weight model, lands at $1.40/MTok. For enterprises processing millions of tokens daily, this price disparity translates to order-of-magnitude differences in operating cost. The implication is clear: organizations that design for model portability can arbitrage this pricing spread, routing premium tasks to S-tier models and commodity tasks to cost-effective open alternatives.

$ Pricing Comparison — Cost Per Million Input Tokens
ModelInput $/MTokOutput $/MTokTypeCost vs Performance
Gemini Flash-Lite $0.10$0.40Closed
DeepSeek V4 $0.14$0.28Open
GPT-5.6 Luna $1.00$6.00Closed
GLM-5.2 $1.40$4.40Open MIT
Gemini 3.1 Pro $2.00$12.00Closed
Claude Sonnet 5 $2.00$10.00Closed
Grok 4.5 $2.00$6.00Closed
Opus 4.8 / GPT-5.6 Sol $5.00$25–30Closed
Claude Fable 5 (Mythos) $10.00$50.00Closed
Frontier Pro reasoning (est.) ~$21~$105Closed
Open-source models can be self-hosted for near-zero marginal cost. Qwen 3.6, OLMo 3.1, and Llama 4 are free to use under their respective licenses. Output prices per facts sheet (July 2026); the frontier reasoning tier is an editorial estimate.
D Distance From Best — Gap Analysis (percentage points behind leader)
Shows how far each model is from the benchmark leader. Green = leader (0pt gap). Red = large gap (>15pt).
Benchmark Leader Gemini
3.1 Pro
Claude
Opus 4.8
GPT
5.6
Grok
4.5
GLM-5.2 Kimi
K3
DS
V4
Qwen
3.6
ARC-AGI-2 Gemini 0.0 -8.3 -22.9
GPQA Diamond Gemini 0.0 -3.3 -5.3 -3.1
AIME 2025 GPT/Grok -5.0 0.0 0.0 -3.9 -4.0 -8.0
SWE-bench Mythos 5 -13.5 -6.9 -10.5 -7.5 -8.5
HLE (tools) GLM-5.2 -4.9 0.0 -0.2
HMMT 2025 Grok 4.5 0.0 -13.7
Who Leads Each Domain (Crown Count)
Gemini 3.1 Pro
2
Reasoning, Science
Claude (family)
3
Coding, Arena, Agentic
GPT-5.6
2
Math, Reliability
Grok 4.5
2
Math Competitions, Writing
GLM-5.2
1
GPQA 91.2%, Open #1
Kimi K3
1
Agentic (Open)
DeepSeek V4
1
Cheapest capable
Qwen 3.6
1
Multilingual, Ecosystem
A Architecture & Specifications
Model Organization Parameters Active Params Architecture Context Window Training Hardware License Release
Gemini 3.1 ProGoogle DeepMindUndisclosedUndisclosedDense (est.)~2M tokensTPU v5p/v6Proprietarycurrent
Claude Opus 4.8AnthropicUndisclosedUndisclosedDense (est.)1M tokensAWS Trainium/GPUProprietarycurrent
Claude Sonnet 5AnthropicUndisclosedUndisclosedDense (est.)1M tokensAWS Trainium/GPUProprietaryJun 30 2026
GPT-5.6OpenAIUndisclosedUndisclosedDense + CoT256K tokensNVIDIA H100/H200ProprietaryJun/Jul 2026
Grok 4.5xAIUndisclosedUndisclosedDense (est.)256K+ tokensColossus (H100)ProprietaryJul 2026
GLM-5.2Zhipu AI744B (est.)UndisclosedMoE (est.)1M tokensHuawei AscendMITcurrent
Kimi K3Moonshot AI~2.8T~32B (est.)MoE128K tokensNVIDIAOpen weightJul 16 2026
DeepSeek V4DeepSeek671B (est.)37B (est.)MoE128K tokensNVIDIA H800Open weightApr 2026
Qwen 3.6Alibaba235B22BMoE128K tokensNVIDIA/AlibabaApache 2.02026
Llama 4 ScoutMeta109B17BMoE (16 exp)10M tokensNVIDIA H100Meta License2025
OLMo 3.1AI232B32BDense32K tokensNVIDIAFully Open2025

Strategic Takeaway

The findings below reinforce the central thesis of the convergence analysis: no single model wins across all dimensions, and the gap between open and closed is no longer generational. Enterprise strategy should be built around model portability, not provider lock-in. The organizations that will extract the most value from AI in the coming years are those that can fluidly route between tiers based on task requirements, cost constraints, and latency needs.

! Key Findings — July 2026

No Single Winner

Gemini leads reasoning (GPQA: 94.3%), Anthropic dominates coding (SWE-bench: 88.6% for Opus 4.8, ~95.5% for its Mythos-class specialists) and human preference (#1 Arena), and GPT-5.6 anchors math and factuality. Choosing the "best" model depends entirely on your use case.

Open-Source Reaches Frontier

On GPQA Diamond, GLM-5.2 (91.2%, open MIT) now sits about 3 points behind Gemini 3.1 Pro (94.3%). On Humanity's Last Exam, every model — open and closed — still scores below 51%, and the best open weights sit within a few points of the frontier. The gap is no longer a generation — it's often single digits.

Chinese Dominance in Open-Source

4 of the top 5 open models are Chinese (GLM-5.2, Kimi K3, Qwen 3.6, DeepSeek V4). Qwen alone has 700M+ downloads. A large share of AI startups build on Chinese open-source. The West leads closed models; China leads open models.

Pricing Spread: 100–350x Between Tiers

Cheap capable open models run at $0.10–0.14/MTok input (Gemini Flash-Lite, DeepSeek V4), while a frontier reasoning tier reaches ~$21 — a spread of roughly 150x on input and far wider on output (DeepSeek V4 $0.28 vs Fable 5 $50 ≈ 180x). Cost is rapidly becoming a routing decision, not a differentiator.

Agentic = New Battleground

Kimi K3 executes hundreds of sequential tool calls autonomously. Grok 4.5 uses a multi-agent collaboration system. Claude Sonnet 5 has native agentic capabilities. Models are now judged on autonomous task completion, not just Q&A.

The February 2026 Release Wave

February 2026 saw a cluster of major releases: Claude Opus 4.6, GPT-5.3 Codex, Gemini 3.1 Pro, GLM-5, Grok 4.1, and Qwen 3.5. The mid-2026 wave kept the pace: DeepSeek V4 (Apr), Claude Sonnet 5 (Jun 30) and Opus 4.8, the GPT-5.6 family (Jun/Jul), Grok 4.5 (Jul), and Kimi K3 (Jul 16). Frontier advancement is accelerating, not slowing.

Sources: Vellum Benchmarks, OpenAI Official Blog, Google DeepMind Blog, LMSYS Chatbot Arena, Artificial Analysis, VentureBeat, HuggingFace, ArXiv, Interconnects, Allen AI, Moonshot AI, DeepSeek, Alibaba Qwen Blog · Data as of July 18, 2026 (original February 2026 edition). Benchmarks are reported from official sources where available; July leaders not yet published are flagged as editorial estimates. "—" indicates no published score. Composite scores are editorial estimates based on available benchmark data. · Created for research purposes.