Compare LLM Models: Pricing, Context, Capabilities & Benchmarks
Pick two or three models and compare the main fields on one page. Compare OpenAI, Anthropic, Google, Mistral, DeepSeek, and other LLM API pricing, context windows, capabilities, and benchmark results side by side to find the best model for your workload.
Data built Sep 19, 2026. Verify with providers before decisions.
Up to 3 models can be compared at once.
Already know your token counts? Estimate cost directly in the calculator.
Start faster
Use a preset if you have not picked models yet
These presets reuse existing article and use-case scenarios. They only preselect models; the comparison table still uses the current database fields.
By workload
Compact model article
Start with GPT-5.4 Mini, Gemini 3.1 Flash Lite, and Claude Haiku 4.5.
Low-cost chat shortlist
Open the lowest-cost text-output chat candidates from the current chatbot guide scenario.
RAG context shortlist
Open the context-ready low-cost candidates from the current RAG guide scenario.
Summarization shortlist
Open the low-cost input-heavy candidates from the current summarization guide scenario.
Coding-agent shortlist
Open the context-ready low-cost candidates from the current coding-agent guide scenario.
Cheapest output models
Models with the lowest output token price. Best for report writing, story generation, long-form content, and document creation.
By price tier
Budget tier models
Under $0.30 / 1M tokens. Fast and cheap for classification, extraction, and simple Q&A.
Mid-tier models
$0.30 to $2.00 / 1M tokens. Balanced capability for RAG, summarization, and customer chat.
Premium tier models
$2.00+ / 1M tokens. Frontier models for complex reasoning, coding, and analysis.
Reasoning tier models
Chain-of-thought models for math, logic, multi-step planning, and self-critique chains.
By feature
Agentic AI models
Function-calling models sorted by combined price. Top picks for tool-use and multi-step agent workflows.
Cache reuse models
Models with cached input pricing sorted by combined price. Best value when repeated context allows cache hits.
Batch processing models
Models with batch API pricing sorted by combined batch price. Cheapest paths for async, high-volume jobs.
Lowest output-cost ratio
Models where output is cheapest relative to input. Best starting point for generation-heavy workloads. See the output-vs-input pricing guide.
Model routing cascade
Budget, mid, and premium models — the 60/30/10 rule for agentic routing: 60% budget, 30% mid, 10% premium.
Multimodal models
Vision-capable models sorted by combined price. Compare GPT-4o, Gemini, Claude, and other models that accept image inputs.
Embedding models
Lowest-cost embedding models for vector search, RAG, and semantic similarity. Compare per-token pricing across providers.
By provider
Open-weight providers
Three lowest-cost provider representatives in this preset: Deepinfra, Lambda AI, Nscale. Compare their current input and output prices side by side.
Cheapest by provider
Three lowest-cost representatives among OpenAI, Anthropic, Google, Mistral, and DeepSeek: Mistral, Google, DeepSeek.
By benchmark
Artificial Analysis Coding Index leaders
Top scores within the same Artificial Analysis Coding Index (score) result set. This preset does not compare raw scores across different benchmarks.
Best Artificial Analysis Coding Index value
Highest Artificial Analysis Coding Index (score) score per dollar of output cost, using one comparable benchmark result set only.
Cost estimates use the generated model database last built on Sep 19, 2026. Pricing, lifecycle, and capability fields can be incomplete or provider-specific, so verify production decisions with the official provider.