Best AI models for reasoning

The strongest AI models for hard reasoning, ranked by GPQA Diamond (a graduate-level science benchmark). All are available on anyAInow — switch between them or compare them side by side, pay-as-you-go.

#ModelGPQA Diamond
1GPT-5.6 Sol
OpenAIFlagship·from $50 / 1M
94.1%Try →
2Gemini 3.1 Pro
GoogleFlagship·from $20 / 1M
94.1%Try →
3GPT-5.5
OpenAIFlagship·from $50 / 1M
93.5%Try →
4Kimi K3
Moonshot (Kimi)Flagship·from $30 / 1M
93.5%Try →
5Claude Opus 5
AnthropicFlagship·from $40 / 1M
93.2%Try →
6Grok 4.5
xAIFlagship·from $20 / 1M
93.1%Try →
7MiniMax M3
MiniMaxOpen·from $10 / 1M
92.9%Try →
8Claude Fable 5
AnthropicFlagship·from $60 / 1M
92.6%Try →
9GPT-5.6 Terra
OpenAIBalanced·from $30 / 1M
92.5%Try →
10Qwen 3.7 Max
QwenFlagship·from $10 / 1M
92.3%Try →
11Gemini 3.5 Flash
GoogleBalanced·from $20 / 1M
92.2%Try →
12GPT-5.4
OpenAIBalanced·from $30 / 1M
92%Try →
13Claude Opus 4.8
AnthropicFlagship·from $30 / 1M
92%Try →
14GPT-5.6 Luna
OpenAIFast·from $10 / 1M
91.1%Try →
15Claude Sonnet 5
AnthropicBalanced·from $20 / 1M
91.1%Try →
16Grok 4.20
xAIBalanced·from $10 / 1M
91.1%Try →
17Kimi K2
Moonshot (Kimi)Balanced·from $10 / 1M
91.1%Try →
18Grok 4.3
xAIFlagship·from $10 / 1M
90.1%Try →
19Qwen 3.7 Plus
QwenOpen·from $10 / 1M
90%Try →
20Grok Build
xAIFast·from $10 / 1M
89.5%Try →
21GLM 5.2
Z.ai (GLM)Open·from $10 / 1M
89.5%Try →
22DeepSeek V4 Flash
DeepSeekFast·from $5 / 1M
89.4%Try →
23DeepSeek V4 Pro
DeepSeekOpen·from $10 / 1M
88.8%Try →
24Qwen 3.6 Plus
QwenOpen·from $10 / 1M
88.2%Try →
25GPT-5.4 mini
OpenAIFast·from $10 / 1M
87.5%Try →
26Claude Sonnet 4.6
AnthropicBalanced·from $30 / 1M
87.5%Try →
27MiniMax M2.7
MiniMaxOpen·from $10 / 1M
87.4%Try →
28MiMo V2.5
Xiaomi (MiMo)Open·from $5 / 1M
86.6%Try →
29Gemini 2.5 Pro
GoogleFlagship·from $20 / 1M
84.4%Try →
30Gemini 3.1 Flash Lite
GoogleFast·from $10 / 1M
82.2%Try →
31DeepSeek R1
DeepSeekReasoning·from $10 / 1M
81.3%Try →
32Mistral Small 4
MistralFast·from $5 / 1M
76.9%Try →
33Mistral Medium 3.5
MistralBalanced·from $20 / 1M
74.8%Try →
34Mistral Large 3
MistralFlagship·from $10 / 1M
68%Try →
35GPT-5 nano
OpenAIFree·Free
67.6%Try →
36Claude Haiku 4.5
AnthropicFast·from $10 / 1M
67.2%Try →
37Llama 4 Maverick
Meta LlamaBalanced·from $10 / 1M
67.1%Try →
38Llama 4 Scout
Meta LlamaOpen·from $5 / 1M
58.7%Try →
39Phi-4
MicrosoftOpen·from $10 / 1M
57.5%Try →
40Nova Pro
AmazonBalanced·from $10 / 1M
49.9%Try →
41Llama 3.3 70B
Meta LlamaFree·Free
49.8%Try →
42Nova Lite
AmazonFast·from $5 / 1M
43.3%Try →
43Gemma 3 27B
GoogleOpen·from $10 / 1M
42.8%Try →

Benchmarks via artificialanalysis.ai, as of 13 Aug 2026. Figures are publisher-reported where not independently verified. Prices shown are the anyAInow pay-as-you-go floor (per 1M tokens); free-tier models deduct nothing.

More rankings