You can combine filters. Only models matching every selection will remain.
Suggested starting points
Three options for different kinds of work. This is not a best-to-worst ranking.
Claude Fable 5.1
top pick
Anthropic
Summary
coding pick backed by current AA.
Claude Fable 5.1 has IQ 53.4 and input $10/M in the sources. Consider it for coding, AI agents; for mass-volume, real-time-latency, run a second benchmark before rollout.
DeepSeek V4.1 Flash has IQ 39.5 and input $0.3/M in the sources. Consider it for large batches, fast replies; for deep-frontier-reasoning, top-coding, run a second benchmark before rollout.
+Good forlarge batches · fast replies · self-hosted stack
−Not suited todeep frontier reasoning · top coding
GLM-5.3 has IQ 44.8 and input $1.4/M in the sources. Consider it for self-hosted stack, sensitive deployments; for premium-agents, top-coding, run a second benchmark before rollout.
+Good forself-hosted stack · sensitive deployments · large batches
A concise overview. Detailed numbers, sources and reasoning live on each model page.
No model matches this combination. Try fewer filters.
ModelSummaryGood fit / poor fitData
Claude Fable 5.1
Anthropic
Summary
coding pick backed by current AA.
Claude Fable 5.1 has IQ 53.4 and input $10/M in the sources. Consider it for coding, AI agents; for mass-volume, real-time-latency, run a second benchmark before rollout.
GPT-6 Astra has IQ 52.7 and input $10/M in the sources. Consider it for coding, AI agents; for mass-volume, self-hosted stack, run a second benchmark before rollout.
+Good forcoding · AI agents · vision and multimodal
−Not suited tomass volume · self-hosted stack · strict latency
Claude Opus 5 has IQ 50.8 and input $5/M in the sources. Consider it for coding, AI agents; for mass-volume, real-time-latency, run a second benchmark before rollout.
GPT-5.6 Sol has IQ 47 and input $4/M in the sources. Consider it for coding, AI agents; for self-hosted stack, real-time-latency, run a second benchmark before rollout.
Qwen3.8 Max has IQ 45.4 and input $2/M in the sources. Consider it for multilingual content, coding; for self-hosted stack, premium-agents, run a second benchmark before rollout.
+Good formultilingual content · coding · large batches
GLM-5.3 has IQ 44.8 and input $1.4/M in the sources. Consider it for self-hosted stack, sensitive deployments; for premium-agents, top-coding, run a second benchmark before rollout.
+Good forself-hosted stack · sensitive deployments · large batches
Grok 4.6 has IQ 44.3 and input $2/M in the sources. Consider it for coding, fast replies; for self-hosted stack, sensitive deployments, run a second benchmark before rollout.
+Good forcoding · fast replies · document extraction
DeepSeek V4.1 Flash has IQ 39.5 and input $0.3/M in the sources. Consider it for large batches, fast replies; for deep-frontier-reasoning, top-coding, run a second benchmark before rollout.
+Good forlarge batches · fast replies · self-hosted stack
−Not suited todeep frontier reasoning · top coding
Gemini 3.6 Flash has IQ 34 and input $0.75/M in the sources. Consider it for large batches, company knowledge; for deep-coding, real-time-latency, run a second benchmark before rollout.
+Good forlarge batches · company knowledge · fast replies
DeepSeek V4 Pro 0813 has IQ 36 and input $1.32/M in the sources. Consider it for large batches, coding; for enterprise-governance, premium-agents, run a second benchmark before rollout.
Command A Plus has IQ 13.1 and input $0/M in the sources. Consider it for company knowledge, document extraction; for top-coding, deep-frontier-reasoning, run a second benchmark before rollout.
Llama 4 Maverick has IQ 10 and input $0.26/M in the sources. Consider it for self-hosted stack, sensitive deployments; for managed-api-comfort, deep-frontier-reasoning, run a second benchmark before rollout.
+Good forself-hosted stack · sensitive deployments · company knowledge
−Not suited tomanaged API comfort · deep frontier reasoning
Kimi K3 has IQ 43.6 and input $3/M in the sources. Consider it for coding, large batches; for enterprise-governance, tool-use, run a second benchmark before rollout.
+Good forcoding · large batches · company knowledge
−Not suited tocompany controls and audit · reliable tool use
Claude Sonnet 5 has IQ 38.2 and input $2/M in the sources. Consider it for coding, AI agents; for self-hosted stack, mass-volume, run a second benchmark before rollout.
Gemini 3.1 Pro Preview has IQ 29.7 and input $2/M in the sources. Consider it for company knowledge, vision and multimodal; for real-time-latency, self-hosted stack, run a second benchmark before rollout.
+Good forcompany knowledge · vision and multimodal · multilingual content
Mistral Medium 3.5 has IQ 14.2 and input $1.5/M in the sources. Consider it for sensitive deployments, company knowledge; for deep-frontier-reasoning, top-coding, run a second benchmark before rollout.
+Good forsensitive deployments · company knowledge · document extraction
−Not suited todeep frontier reasoning · top coding
A curated model selection updated daily. Prices and the Intelligence Index are checked against fresh Artificial Analysis data; additional benchmarks identify their own sources. A change in benchmark methodology can change a score without a change to the model, so older and newer index values may not be directly comparable. If validation fails, the last verified snapshot keeps its original date.