Best AI models for reasoning
The strongest AI models for hard reasoning, ranked by GPQA Diamond (a graduate-level science benchmark). All are available on anyAInow — switch between them or compare them side by side, pay-as-you-go.
| # | Model | GPQA Diamond | |
|---|---|---|---|
| 1 | GPT-5.6 Sol OpenAIFlagship·from $50 / 1M | 94.1% | Try → |
| 2 | Gemini 3.1 Pro GoogleFlagship·from $20 / 1M | 94.1% | Try → |
| 3 | GPT-5.5 OpenAIFlagship·from $50 / 1M | 93.5% | Try → |
| 4 | Kimi K3 Moonshot (Kimi)Flagship·from $30 / 1M | 93.5% | Try → |
| 5 | Claude Opus 5 AnthropicFlagship·from $40 / 1M | 93.2% | Try → |
| 6 | Grok 4.5 xAIFlagship·from $20 / 1M | 93.1% | Try → |
| 7 | MiniMax M3 MiniMaxOpen·from $10 / 1M | 92.9% | Try → |
| 8 | Claude Fable 5 AnthropicFlagship·from $60 / 1M | 92.6% | Try → |
| 9 | GPT-5.6 Terra OpenAIBalanced·from $30 / 1M | 92.5% | Try → |
| 10 | Qwen 3.7 Max QwenFlagship·from $10 / 1M | 92.3% | Try → |
| 11 | Gemini 3.5 Flash GoogleBalanced·from $20 / 1M | 92.2% | Try → |
| 12 | GPT-5.4 OpenAIBalanced·from $30 / 1M | 92% | Try → |
| 13 | Claude Opus 4.8 AnthropicFlagship·from $30 / 1M | 92% | Try → |
| 14 | GPT-5.6 Luna OpenAIFast·from $10 / 1M | 91.1% | Try → |
| 15 | Claude Sonnet 5 AnthropicBalanced·from $20 / 1M | 91.1% | Try → |
| 16 | Grok 4.20 xAIBalanced·from $10 / 1M | 91.1% | Try → |
| 17 | Kimi K2 Moonshot (Kimi)Balanced·from $10 / 1M | 91.1% | Try → |
| 18 | Grok 4.3 xAIFlagship·from $10 / 1M | 90.1% | Try → |
| 19 | Qwen 3.7 Plus QwenOpen·from $10 / 1M | 90% | Try → |
| 20 | Grok Build xAIFast·from $10 / 1M | 89.5% | Try → |
| 21 | GLM 5.2 Z.ai (GLM)Open·from $10 / 1M | 89.5% | Try → |
| 22 | DeepSeek V4 Flash DeepSeekFast·from $5 / 1M | 89.4% | Try → |
| 23 | DeepSeek V4 Pro DeepSeekOpen·from $10 / 1M | 88.8% | Try → |
| 24 | Qwen 3.6 Plus QwenOpen·from $10 / 1M | 88.2% | Try → |
| 25 | GPT-5.4 mini OpenAIFast·from $10 / 1M | 87.5% | Try → |
| 26 | Claude Sonnet 4.6 AnthropicBalanced·from $30 / 1M | 87.5% | Try → |
| 27 | MiniMax M2.7 MiniMaxOpen·from $10 / 1M | 87.4% | Try → |
| 28 | MiMo V2.5 Xiaomi (MiMo)Open·from $5 / 1M | 86.6% | Try → |
| 29 | Gemini 2.5 Pro GoogleFlagship·from $20 / 1M | 84.4% | Try → |
| 30 | Gemini 3.1 Flash Lite GoogleFast·from $10 / 1M | 82.2% | Try → |
| 31 | DeepSeek R1 DeepSeekReasoning·from $10 / 1M | 81.3% | Try → |
| 32 | Mistral Small 4 MistralFast·from $5 / 1M | 76.9% | Try → |
| 33 | Mistral Medium 3.5 MistralBalanced·from $20 / 1M | 74.8% | Try → |
| 34 | Mistral Large 3 MistralFlagship·from $10 / 1M | 68% | Try → |
| 35 | GPT-5 nano OpenAIFree·Free | 67.6% | Try → |
| 36 | Claude Haiku 4.5 AnthropicFast·from $10 / 1M | 67.2% | Try → |
| 37 | Llama 4 Maverick Meta LlamaBalanced·from $10 / 1M | 67.1% | Try → |
| 38 | Llama 4 Scout Meta LlamaOpen·from $5 / 1M | 58.7% | Try → |
| 39 | Phi-4 MicrosoftOpen·from $10 / 1M | 57.5% | Try → |
| 40 | Nova Pro AmazonBalanced·from $10 / 1M | 49.9% | Try → |
| 41 | Llama 3.3 70B Meta LlamaFree·Free | 49.8% | Try → |
| 42 | Nova Lite AmazonFast·from $5 / 1M | 43.3% | Try → |
| 43 | Gemma 3 27B GoogleOpen·from $10 / 1M | 42.8% | Try → |
Benchmarks via artificialanalysis.ai, as of 13 Aug 2026. Figures are publisher-reported where not independently verified. Prices shown are the anyAInow pay-as-you-go floor (per 1M tokens); free-tier models deduct nothing.