Lilith Lilith.

LLM compare

Model comparison

A practical overview by use case, price and limitations. Apply filters, then open the model that fits the work at hand.

Updated 2026-09-23 Next review 2026-09-24 Change history →
Filter models
16 models
Use case
Budget
Category

You can combine filters. Only models matching every selection will remain.

Suggested starting points

Three options for different kinds of work. This is not a best-to-worst ranking.

Claude Fable 5.1

top pick
Anthropic

Summary

coding pick backed by current AA.

Claude Fable 5.1 has IQ 53.4 and input $10/M in the sources. Consider it for coding, AI agents; for mass-volume, real-time-latency, run a second benchmark before rollout.

+Good forcoding · AI agents · company knowledge
Not suited tomass volume · strict latency
53.4intelligence index $10.0input / 1M $50.0output / 1M

premium, when mistakes hurt · market frontier · verified by external data

market frontier premium, when mistakes hurt codingAI agents
Model details

DeepSeek V4.1 Flash

top pick
DeepSeek

Summary

batch pick backed by current AA.

DeepSeek V4.1 Flash has IQ 39.5 and input $0.3/M in the sources. Consider it for large batches, fast replies; for deep-frontier-reasoning, top-coding, run a second benchmark before rollout.

+Good forlarge batches · fast replies · self-hosted stack
Not suited todeep frontier reasoning · top coding
39.5intelligence index $0.3input / 1M $1.2output / 1M

open / self-run · specialist · verified by external data

specialist open / self-run large batchesfast replies
Model details

GLM-5.3

top pick
Z.AI/Zhipu

Summary

self-hosted pick backed by current AA.

GLM-5.3 has IQ 44.8 and input $1.4/M in the sources. Consider it for self-hosted stack, sensitive deployments; for premium-agents, top-coding, run a second benchmark before rollout.

+Good forself-hosted stack · sensitive deployments · large batches
Not suited topremium agents · top coding
44.8intelligence index $1.4input / 1M $4.4output / 1M

open / self-run · specialist · verified by external data

specialist open / self-run self-hosted stacksensitive deployments
Model details

Models

A concise overview. Detailed numbers, sources and reasoning live on each model page.

Claude Fable 5.1

Anthropic

Summary

coding pick backed by current AA.

Claude Fable 5.1 has IQ 53.4 and input $10/M in the sources. Consider it for coding, AI agents; for mass-volume, real-time-latency, run a second benchmark before rollout.

+Good forcoding · AI agents · company knowledge
Not suited tomass volume · strict latency
53.4intelligence index $10.0input / 1M $50.0output / 1M

premium, when mistakes hurt · market frontier · verified by external data

market frontier premium, when mistakes hurt codingAI agents
Model details

GPT-6 Astra

OpenAI

Summary

coding pick backed by current AA.

GPT-6 Astra has IQ 52.7 and input $10/M in the sources. Consider it for coding, AI agents; for mass-volume, self-hosted stack, run a second benchmark before rollout.

+Good forcoding · AI agents · vision and multimodal
Not suited tomass volume · self-hosted stack · strict latency
52.7intelligence index $10.0input / 1M $50.0output / 1M

premium, when mistakes hurt · market frontier · verified by external data

market frontier premium, when mistakes hurt codingAI agents
Model details

Claude Opus 5

Anthropic

Summary

coding pick backed by current AA.

Claude Opus 5 has IQ 50.8 and input $5/M in the sources. Consider it for coding, AI agents; for mass-volume, real-time-latency, run a second benchmark before rollout.

+Good forcoding · AI agents · company knowledge
Not suited tomass volume · strict latency
50.8intelligence index $5.0input / 1M $25.0output / 1M

mid-budget · market frontier · verified by external data

market frontier mid-budget codingAI agents
Model details

GPT-5.6 Sol

OpenAI

Summary

coding pick backed by current AA.

GPT-5.6 Sol has IQ 47 and input $4/M in the sources. Consider it for coding, AI agents; for self-hosted stack, real-time-latency, run a second benchmark before rollout.

+Good forcoding · AI agents · document extraction
Not suited toself-hosted stack · strict latency
47.0intelligence index $4.0input / 1M $20.0output / 1M

mid-budget · market frontier · verified by external data

market frontier mid-budget codingAI agents
Model details

Qwen3.8 Max

Alibaba

Summary

multilingual pick backed by current AA.

Qwen3.8 Max has IQ 45.4 and input $2/M in the sources. Consider it for multilingual content, coding; for self-hosted stack, premium-agents, run a second benchmark before rollout.

+Good formultilingual content · coding · large batches
Not suited toself-hosted stack · premium agents
45.4intelligence index $2.0input / 1M $6.0output / 1M

mid-budget · market frontier · verified by external data

market frontier mid-budget multilingual contentcoding
Model details

GLM-5.3

Z.AI/Zhipu

Summary

self-hosted pick backed by current AA.

GLM-5.3 has IQ 44.8 and input $1.4/M in the sources. Consider it for self-hosted stack, sensitive deployments; for premium-agents, top-coding, run a second benchmark before rollout.

+Good forself-hosted stack · sensitive deployments · large batches
Not suited topremium agents · top coding
44.8intelligence index $1.4input / 1M $4.4output / 1M

open / self-run · specialist · verified by external data

specialist open / self-run self-hosted stacksensitive deployments
Model details

Grok 4.6

xAI

Summary

coding pick backed by current AA.

Grok 4.6 has IQ 44.3 and input $2/M in the sources. Consider it for coding, fast replies; for self-hosted stack, sensitive deployments, run a second benchmark before rollout.

+Good forcoding · fast replies · document extraction
Not suited toself-hosted stack · sensitive deployments
44.3intelligence index $2.0input / 1M $6.0output / 1M

mid-budget · specialist · verified by external data

specialist mid-budget codingfast replies
Model details

DeepSeek V4.1 Flash

DeepSeek

Summary

batch pick backed by current AA.

DeepSeek V4.1 Flash has IQ 39.5 and input $0.3/M in the sources. Consider it for large batches, fast replies; for deep-frontier-reasoning, top-coding, run a second benchmark before rollout.

+Good forlarge batches · fast replies · self-hosted stack
Not suited todeep frontier reasoning · top coding
39.5intelligence index $0.3input / 1M $1.2output / 1M

open / self-run · specialist · verified by external data

specialist open / self-run large batchesfast replies
Model details

Gemini 3.6 Flash

Google

Summary

batch pick backed by current AA.

Gemini 3.6 Flash has IQ 34 and input $0.75/M in the sources. Consider it for large batches, company knowledge; for deep-coding, real-time-latency, run a second benchmark before rollout.

+Good forlarge batches · company knowledge · fast replies
Not suited tohard coding work · strict latency
34.0intelligence index $0.75input / 1M $3.75output / 1M

cheap at volume · specialist · verified by external data

specialist cheap at volume large batchescompany knowledge
Model details

DeepSeek V4 Pro 0813

DeepSeek

Summary

batch pick backed by current AA.

DeepSeek V4 Pro 0813 has IQ 36 and input $1.32/M in the sources. Consider it for large batches, coding; for enterprise-governance, premium-agents, run a second benchmark before rollout.

+Good forlarge batches · coding · self-hosted stack
Not suited tocompany controls and audit · premium agents
36.0intelligence index $1.32input / 1M $3.96output / 1M

open / self-run · specialist · verified by external data

specialist open / self-run large batchescoding
Model details

Command A+

Cohere

Summary

rag pick backed by current AA.

Command A Plus has IQ 13.1 and input $0/M in the sources. Consider it for company knowledge, document extraction; for top-coding, deep-frontier-reasoning, run a second benchmark before rollout.

+Good forcompany knowledge · document extraction · sensitive deployments
Not suited totop coding · deep frontier reasoning
13.1intelligence index $0.0input / 1M $0.0output / 1M

cheap at volume · specialist · verified by external data

specialist cheap at volume company knowledgedocument extraction
Model details

Llama 4 Maverick

Meta

Summary

self-hosted pick backed by current AA.

Llama 4 Maverick has IQ 10 and input $0.26/M in the sources. Consider it for self-hosted stack, sensitive deployments; for managed-api-comfort, deep-frontier-reasoning, run a second benchmark before rollout.

+Good forself-hosted stack · sensitive deployments · company knowledge
Not suited tomanaged API comfort · deep frontier reasoning
10.0intelligence index $0.26input / 1M $0.91output / 1M

open / self-run · specialist · verified by external data

specialist open / self-run self-hosted stacksensitive deployments
Model details

Kimi K3

Moonshot

Summary

coding pick backed by current AA.

Kimi K3 has IQ 43.6 and input $3/M in the sources. Consider it for coding, large batches; for enterprise-governance, tool-use, run a second benchmark before rollout.

+Good forcoding · large batches · company knowledge
Not suited tocompany controls and audit · reliable tool use
43.6intelligence index $3.0input / 1M $15.0output / 1M

mid-budget · specialist · verified by external data

specialist mid-budget codinglarge batches
Model details

Claude Sonnet 5

Anthropic

Summary

coding pick backed by current AA.

Claude Sonnet 5 has IQ 38.2 and input $2/M in the sources. Consider it for coding, AI agents; for self-hosted stack, mass-volume, run a second benchmark before rollout.

+Good forcoding · AI agents · company knowledge
Not suited toself-hosted stack · mass volume
38.2intelligence index $2.0input / 1M $10.0output / 1M

mid-budget · specialist · verified by external data

specialist mid-budget codingAI agents
Model details

Gemini 3.1 Pro Preview

Google

Summary

rag pick backed by current AA.

Gemini 3.1 Pro Preview has IQ 29.7 and input $2/M in the sources. Consider it for company knowledge, vision and multimodal; for real-time-latency, self-hosted stack, run a second benchmark before rollout.

+Good forcompany knowledge · vision and multimodal · multilingual content
Not suited tostrict latency · self-hosted stack
29.7intelligence index $2.0input / 1M $12.0output / 1M

mid-budget · specialist · verified by external data

specialist mid-budget company knowledgevision and multimodal
Model details

Mistral Medium 3.5

Mistral AI

Summary

compliance pick backed by current AA.

Mistral Medium 3.5 has IQ 14.2 and input $1.5/M in the sources. Consider it for sensitive deployments, company knowledge; for deep-frontier-reasoning, top-coding, run a second benchmark before rollout.

+Good forsensitive deployments · company knowledge · document extraction
Not suited todeep frontier reasoning · top coding
14.2intelligence index $1.5input / 1M $7.5output / 1M

mid-budget · specialist · verified by external data

specialist mid-budget sensitive deploymentscompany knowledge
Model details
How the data is assembled

A curated model selection updated daily. Prices and the Intelligence Index are checked against fresh Artificial Analysis data; additional benchmarks identify their own sources. A change in benchmark methodology can change a score without a change to the model, so older and newer index values may not be directly comparable. If validation fails, the last verified snapshot keeps its original date.