LLM Comparison Matrix
Ratings are directional, not absolute. Updated August 7, 2026 with
latest releases and pricing changes. Use 1-5 ranking to compare faster,
then sort by what matters most for your product.
Latest matrix includes: Claude Opus 4.9 (new frontier), GPT-5.7
(1.2M context), Gemini 3.6 Pro (stable), Claude Opus 4.8, Claude
Sonnet 5 (speed/intelligence), GPT-5.6 Sol, Gemini 3.6 Flash,
Llama 4.1 (multilingual MoE), DeepSeek V4.1-DSpark (920B),
Qwen3.8-2.4T-A95B (262K–1M context), Kimi K3 (2.8T), GLM-5.3 (753B), and Mistral Medium 3.5.
Sort by
Overall Score
Reasoning
Coding
Cost Efficiency
Latency
Context Quality
Deployment Control
Reset
Model
Overall
Reasoning
Coding
Cost Efficiency
Latency
Context Quality
Deployment Control
Claude Opus 4.9
4.7/5 ★★★★★
5/5 ★★★★★
5/5 ★★★★★
2/5 ★★☆☆☆
3/5 ★★★☆☆
5/5 ★★★★★
2/5 ★★☆☆☆
GPT-5.7
4.6/5 ★★★★★
5/5 ★★★★★
5/5 ★★★★★
2/5 ★★☆☆☆
3/5 ★★★☆☆
5/5 ★★★★★
2/5 ★★☆☆☆
Gemini 3.6 Pro
4.5/5 ★★★★★
5/5 ★★★★★
5/5 ★★★★★
3/5 ★★★☆☆
3/5 ★★★☆☆
5/5 ★★★★★
2/5 ★★☆☆☆
Claude Opus 4.8
4.5/5 ★★★★★
5/5 ★★★★★
5/5 ★★★★★
3/5 ★★★☆☆
3/5 ★★★☆☆
5/5 ★★★★★
2/5 ★★☆☆☆
GPT-5.5-Pro
4.3/5 ★★★★☆
5/5 ★★★★★
5/5 ★★★★★
2/5 ★★☆☆☆
2/5 ★★☆☆☆
5/5 ★★★★★
2/5 ★★☆☆☆
GPT-5.5
4.2/5 ★★★★☆
5/5 ★★★★★
5/5 ★★★★★
3/5 ★★★☆☆
3/5 ★★★☆☆
5/5 ★★★★★
2/5 ★★☆☆☆
GPT-5.4 mini
4.0/5 ★★★★☆
4/5 ★★★★☆
5/5 ★★★★★
4/5 ★★★★☆
4/5 ★★★★☆
4/5 ★★★★☆
2/5 ★★☆☆☆
Claude Sonnet 4.6
4.0/5 ★★★★☆
5/5 ★★★★★
5/5 ★★★★★
3/5 ★★★☆☆
4/5 ★★★★☆
5/5 ★★★★★
2/5 ★★☆☆☆
Claude Haiku 4.5
3.5/5 ★★★★☆
4/5 ★★★★☆
3/5 ★★★☆☆
4/5 ★★★★☆
5/5 ★★★★★
3/5 ★★★☆☆
2/5 ★★☆☆☆
Gemini 3.5 Flash
3.8/5 ★★★★☆
4/5 ★★★★☆
5/5 ★★★★★
4/5 ★★★★☆
5/5 ★★★★★
4/5 ★★★★☆
2/5 ★★☆☆☆
Gemma 4 31B
4.1/5 ★★★★☆
5/5 ★★★★★
4/5 ★★★★☆
5/5 ★★★★★
4/5 ★★★★☆
4/5 ★★★★☆
5/5 ★★★★★
Llama 4 Maverick
4.3/5 ★★★★☆
5/5 ★★★★★
4/5 ★★★★☆
4/5 ★★★★☆
4/5 ★★★★☆
4/5 ★★★★☆
5/5 ★★★★★
Llama 4 Scout
4.0/5 ★★★★☆
4/5 ★★★★☆
4/5 ★★★★☆
5/5 ★★★★★
5/5 ★★★★★
3/5 ★★★☆☆
5/5 ★★★★★
Mistral Medium 3.5
3.8/5 ★★★★☆
4/5 ★★★★☆
5/5 ★★★★★
4/5 ★★★★☆
4/5 ★★★★☆
3/5 ★★★☆☆
4/5 ★★★★☆
Qwen3.5-35B-A3B
4.0/5 ★★★★☆
4/5 ★★★★☆
5/5 ★★★★★
5/5 ★★★★★
4/5 ★★★★☆
4/5 ★★★★☆
4/5 ★★★★☆
DeepSeek V4-Pro
4.0/5 ★★★★☆
5/5 ★★★★★
4/5 ★★★★☆
4/5 ★★★★☆
3/5 ★★★☆☆
4/5 ★★★★☆
4/5 ★★★★☆
Llama 4.1
4.4/5 ★★★★☆
5/5 ★★★★★
4/5 ★★★★☆
4/5 ★★★★☆
4/5 ★★★★☆
4/5 ★★★★☆
5/5 ★★★★★
DeepSeek V4.1-DSpark
4.2/5 ★★★★☆
5/5 ★★★★★
5/5 ★★★★★
4/5 ★★★★☆
3/5 ★★★☆☆
4/5 ★★★★☆
4/5 ★★★★☆
Qwen3.8-2.4T-A95B
4.1/5 ★★★★☆
4/5 ★★★★☆
5/5 ★★★★★
5/5 ★★★★★
4/5 ★★★★☆
5/5 ★★★★★
4/5 ★★★★☆
Score key: 5 = excellent, 4 = strong, 3 = medium, 2 = low.
Interpreting the Matrix
For customer-facing quality, prioritize reasoning + context quality.
For internal automation at scale, prioritize cost efficiency +
latency.