The definitive guide to AI API costs. 88 models, 10 providers, one page. Updated July 5, 2026.
Ranked by blended cost (75% input, 25% output — typical for most applications). Prices per 1 million tokens.
| # | Model | Provider | Input $/1M | Output $/1M | Blended $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemini 2.0 Flash Lite | $0.075 | $0.30 | $0.13 | 1M | |
| 2 | Llama 3.1 8B | Together | $0.10 | $0.10 | $0.10 | 128K |
| 3 | Gemini 2.5 Flash-Lite | $0.10 | $0.40 | $0.18 | 1M | |
| 4 | Mistral Small 4 | Mistral | $0.10 | $0.30 | $0.15 | 128K |
| 5 | DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | $0.18 | 1M |
| 6 | GPT-4o mini | OpenAI | $0.15 | $0.60 | $0.26 | 128K |
| 7 | GPT-oss 20B | OpenAI | $0.08 | $0.35 | $0.15 | 128K |
| 8 | Llama 4 Scout | Together | $0.18 | $0.59 | $0.28 | 1M |
| 9 | GPT-5.4 nano | OpenAI | $0.20 | $1.25 | $0.46 | 400K |
| 10 | DeepSeek V3.2 | DeepSeek | $0.23 | $0.34 | $0.26 | 128K |
The cheapest AI API is not always the best value. Gemini 2.0 Flash Lite is the cheapest per token, but Llama 3.1 8B has the lowest output cost ($0.10/1M). For output-heavy applications (chatbots, code generation), Llama 3.1 8B or DeepSeek V4 Flash may be more cost-effective despite higher input prices.
Complete pricing for every model tracked by APIpulse. Sort by any column to find the best option for your use case.
| Model | Provider | Tier | Input $/1M | Output $/1M | Context | Status |
|---|---|---|---|---|---|---|
| GPT-5.5 | OpenAI | Premium | $5.00 | $30.00 | 1.05M | Active |
| GPT-5.5 Pro | OpenAI | Premium | $30.00 | $180.00 | 1.05M | Active |
| GPT-5.3 Codex | OpenAI | Mid | $1.75 | $14.00 | 400K | Active |
| GPT-5 | OpenAI | Premium | $1.25 | $10.00 | 272K | Active |
| GPT-5 mini | OpenAI | Budget | $0.25 | $2.00 | 272K | Active |
| GPT-oss 120B | OpenAI | Budget | $0.15 | $0.60 | 128K | Active |
| GPT-oss 20B | OpenAI | Budget | $0.08 | $0.35 | 128K | Active |
| GPT-4o | OpenAI | Mid | $2.50 | $10.00 | 128K | Active |
| GPT-4o mini | OpenAI | Budget | $0.15 | $0.60 | 128K | Active |
| GPT-5.4 | OpenAI | Mid | $2.50 | $15.00 | 400K | Active |
| GPT-5.4 mini | OpenAI | Budget | $0.75 | $4.50 | 400K | Active |
| GPT-5.4 nano | OpenAI | Budget | $0.20 | $1.25 | 400K | Active |
| GPT-5.4 Pro | OpenAI | Premium | $30.00 | $180.00 | 400K | Active |
| Claude Opus 4.8 | Anthropic | Premium | $5.00 | $25.00 | 1M | Active |
| Claude Opus 4.7 | Anthropic | Premium | $5.00 | $25.00 | 1M | Active |
| Claude 4 Opus | Anthropic | Premium | $15.00 | $75.00 | 200K | Deprecated |
| Claude Sonnet 5 | Anthropic | Mid | $3.00 | $15.00 | 1M | Active |
| Claude Sonnet 4.6 | Anthropic | Mid | $3.00 | $15.00 | 1M | Active |
| Claude Sonnet 4 | Anthropic | Mid | $3.00 | $15.00 | 200K | Deprecated |
| Claude Haiku 4.5 | Anthropic | Budget | $1.00 | $5.00 | 200K | Active |
| Claude Fable 5 | Anthropic | Premium | $10.00 | $50.00 | 1M | Active |
| Claude Mythos 5 | Anthropic | Premium | $10.00 | $50.00 | 1M | Active (Limited) |
| Gemini 3.5 Flash | Mid | $1.50 | $9.00 | 1M | Active | |
| Gemini 3.1 Flash-Lite | Budget | $0.25 | $1.50 | 1M | Active | |
| Gemini 3.1 Pro | Mid | $2.00 | $12.00 | 1M | Active | |
| Gemini 3 Flash | Budget | $0.50 | $3.00 | 1M | Active | |
| Gemini 2.5 Pro | Mid | $1.25 | $10.00 | 1M | Active | |
| Gemini 2.5 Flash-Lite | Budget | $0.10 | $0.40 | 1M | Active | |
| Gemini 2.0 Flash | Budget | $0.10 | $0.40 | 1M | Deprecated | |
| Gemini 2.0 Flash Lite | Budget | $0.075 | $0.30 | 1M | Deprecated | |
| DeepSeek V4 Pro | DeepSeek | Budget | $0.435 | $0.87 | 1M | Active |
| DeepSeek V4 Flash | DeepSeek | Budget | $0.14 | $0.28 | 1M | Active |
| DeepSeek V3.2 | DeepSeek | Budget | $0.23 | $0.34 | 128K | Active |
| DeepSeek V3 | DeepSeek | Budget | $0.27 | $1.10 | 128K | Deprecated |
| Mistral Large 3 | Mistral | Budget | $0.50 | $1.50 | 262K | Active |
| Mistral Medium 3.5 | Mistral | Mid | $1.50 | $7.50 | 128K | Active |
| Mistral Small 4 | Mistral | Budget | $0.10 | $0.30 | 128K | Active |
| Command A | Cohere | Mid | $2.50 | $10.00 | 128K | Active |
| Command R+ | Cohere | Mid | $2.50 | $10.00 | 128K | Active |
| Command R | Cohere | Budget | $0.50 | $1.50 | 128K | Active |
| Llama 4 Scout | Together | Budget | $0.18 | $0.59 | 1M | Active |
| Llama 4 Maverick | Together | Budget | $0.27 | $0.85 | 1M | Active |
| Llama 3.3 70B | Together | Mid | $1.04 | $1.04 | 128K | Active |
| Llama 3.1 8B | Together | Budget | $0.10 | $0.10 | 128K | Active |
| Kimi K2.6 | Moonshot | Budget | $0.95 | $4.00 | 256K | Active |
| Kimi K2.7 Code | Moonshot | Budget | $0.96 | $3.97 | 256K | Active |
| Grok 4.3 | xAI | Mid | $1.25 | $2.50 | 1M | Active |
| Grok Build 0.1 | xAI | Budget | $0.30 | $0.50 | 256K | Active |
| Jamba 1.7 Large | AI21 | Mid | $2.00 | $8.00 | 256K | Active |
| Jamba 1.5 Large | AI21 | Mid | $2.00 | $8.00 | 256K | Deprecated |
How each provider stacks up on price, model count, and context window support.
Sub-$0.10/1M input pricing is now standard for budget models. Google leads with Gemini 2.0 Flash Lite at $0.075/1M — that's 7.5 cents per million input tokens. At this price, processing 1 billion tokens costs just $75. This makes AI API costs essentially negligible for most applications.
Claude 4 Opus, Claude Sonnet 4, DeepSeek V3, Gemini 2.0 Flash, Gemini 2.0 Flash Lite, and Jamba 1.5 Large are all deprecated. If you're still using these models, migrate now — pricing and availability are not guaranteed. Check our Model Deprecation Checker for migration paths.
Seven providers now offer 1M+ context windows: OpenAI (1.05M on GPT-5.5), Anthropic (1M on Opus/Sonnet/Haiku), Google (1M on all Gemini), DeepSeek (1M on V4), xAI (1M on Grok 4.3), and Together/Meta (1M on Llama 4). Long-context processing is now table stakes, not a premium feature.
Output tokens cost 3-6× more than input tokens across all providers. For a typical 3:1 input:output ratio, output tokens account for ~43% of your bill despite being only 25% of your tokens. Optimize by keeping responses short, using structured output formats, and avoiding verbose models for simple tasks.
All this data is available as a free, no-auth JSON API. Use it in your own tools, dashboards, or applications.
GET https://www.getapipulse.com/data/pricing.json
No authentication required. 88 models, 10 providers. Updated regularly. Full API docs →
Use our free calculator to see exactly how much you could save by switching models. Or run a full audit to get a personalized recommendation with migration code.