📊 Pricing Report

AI API Pricing Report
July 2026

The definitive guide to AI API costs. 88 models, 10 providers, one page. Updated July 5, 2026.

📅 July 5, 2026 📱 88 models 🌐 10 providers 💾 Free data API
$0.075
Cheapest input price
(Gemini 2.0 Flash Lite)
$180
Most expensive output
(GPT-5.5 Pro, GPT-5.4 Pro)
2,400×
Price gap: cheapest vs most expensive output
67%
Models under $2/1M input

Table of Contents

1. Cheapest AI Models (July 2026) 2. Full Pricing Table — All 67 Models 3. Provider Comparison 4. Premium Models — What You're Paying For 5. Key Insights & Trends 6. Free Pricing API

1. Cheapest AI Models — July 2026

Ranked by blended cost (75% input, 25% output — typical for most applications). Prices per 1 million tokens.

# Model Provider Input $/1M Output $/1M Blended $/1M Context
1 Gemini 2.0 Flash Lite Google $0.075 $0.30 $0.13 1M
2 Llama 3.1 8B Together $0.10 $0.10 $0.10 128K
3 Gemini 2.5 Flash-Lite Google $0.10 $0.40 $0.18 1M
4 Mistral Small 4 Mistral $0.10 $0.30 $0.15 128K
5 DeepSeek V4 Flash DeepSeek $0.14 $0.28 $0.18 1M
6 GPT-4o mini OpenAI $0.15 $0.60 $0.26 128K
7 GPT-oss 20B OpenAI $0.08 $0.35 $0.15 128K
8 Llama 4 Scout Together $0.18 $0.59 $0.28 1M
9 GPT-5.4 nano OpenAI $0.20 $1.25 $0.46 400K
10 DeepSeek V3.2 DeepSeek $0.23 $0.34 $0.26 128K

💡 Key Insight

The cheapest AI API is not always the best value. Gemini 2.0 Flash Lite is the cheapest per token, but Llama 3.1 8B has the lowest output cost ($0.10/1M). For output-heavy applications (chatbots, code generation), Llama 3.1 8B or DeepSeek V4 Flash may be more cost-effective despite higher input prices.

2. Full Pricing Table — All 67 Models

Complete pricing for every model tracked by APIpulse. Sort by any column to find the best option for your use case.

Model Provider Tier Input $/1M Output $/1M Context Status
GPT-5.5 OpenAI Premium $5.00 $30.00 1.05M Active
GPT-5.5 Pro OpenAI Premium $30.00 $180.00 1.05M Active
GPT-5.3 Codex OpenAI Mid $1.75 $14.00 400K Active
GPT-5 OpenAI Premium $1.25 $10.00 272K Active
GPT-5 mini OpenAI Budget $0.25 $2.00 272K Active
GPT-oss 120B OpenAI Budget $0.15 $0.60 128K Active
GPT-oss 20B OpenAI Budget $0.08 $0.35 128K Active
GPT-4o OpenAI Mid $2.50 $10.00 128K Active
GPT-4o mini OpenAI Budget $0.15 $0.60 128K Active
GPT-5.4 OpenAI Mid $2.50 $15.00 400K Active
GPT-5.4 mini OpenAI Budget $0.75 $4.50 400K Active
GPT-5.4 nano OpenAI Budget $0.20 $1.25 400K Active
GPT-5.4 Pro OpenAI Premium $30.00 $180.00 400K Active
Claude Opus 4.8 Anthropic Premium $5.00 $25.00 1M Active
Claude Opus 4.7 Anthropic Premium $5.00 $25.00 1M Active
Claude 4 Opus Anthropic Premium $15.00 $75.00 200K Deprecated
Claude Sonnet 5 Anthropic Mid $3.00 $15.00 1M Active
Claude Sonnet 4.6 Anthropic Mid $3.00 $15.00 1M Active
Claude Sonnet 4 Anthropic Mid $3.00 $15.00 200K Deprecated
Claude Haiku 4.5 Anthropic Budget $1.00 $5.00 200K Active
Claude Fable 5 Anthropic Premium $10.00 $50.00 1M Active
Claude Mythos 5 Anthropic Premium $10.00 $50.00 1M Active (Limited)
Gemini 3.5 Flash Google Mid $1.50 $9.00 1M Active
Gemini 3.1 Flash-Lite Google Budget $0.25 $1.50 1M Active
Gemini 3.1 Pro Google Mid $2.00 $12.00 1M Active
Gemini 3 Flash Google Budget $0.50 $3.00 1M Active
Gemini 2.5 Pro Google Mid $1.25 $10.00 1M Active
Gemini 2.5 Flash-Lite Google Budget $0.10 $0.40 1M Active
Gemini 2.0 Flash Google Budget $0.10 $0.40 1M Deprecated
Gemini 2.0 Flash Lite Google Budget $0.075 $0.30 1M Deprecated
DeepSeek V4 Pro DeepSeek Budget $0.435 $0.87 1M Active
DeepSeek V4 Flash DeepSeek Budget $0.14 $0.28 1M Active
DeepSeek V3.2 DeepSeek Budget $0.23 $0.34 128K Active
DeepSeek V3 DeepSeek Budget $0.27 $1.10 128K Deprecated
Mistral Large 3 Mistral Budget $0.50 $1.50 262K Active
Mistral Medium 3.5 Mistral Mid $1.50 $7.50 128K Active
Mistral Small 4 Mistral Budget $0.10 $0.30 128K Active
Command A Cohere Mid $2.50 $10.00 128K Active
Command R+ Cohere Mid $2.50 $10.00 128K Active
Command R Cohere Budget $0.50 $1.50 128K Active
Llama 4 Scout Together Budget $0.18 $0.59 1M Active
Llama 4 Maverick Together Budget $0.27 $0.85 1M Active
Llama 3.3 70B Together Mid $1.04 $1.04 128K Active
Llama 3.1 8B Together Budget $0.10 $0.10 128K Active
Kimi K2.6 Moonshot Budget $0.95 $4.00 256K Active
Kimi K2.7 Code Moonshot Budget $0.96 $3.97 256K Active
Grok 4.3 xAI Mid $1.25 $2.50 1M Active
Grok Build 0.1 xAI Budget $0.30 $0.50 256K Active
Jamba 1.7 Large AI21 Mid $2.00 $8.00 256K Active
Jamba 1.5 Large AI21 Mid $2.00 $8.00 256K Deprecated

3. Provider Comparison

How each provider stacks up on price, model count, and context window support.

🟢 OpenAI

Models13
Cheapest input$0.08 (GPT-oss 20B)
Most expensive output$180 (GPT-5.5 Pro)
Max context1.05M
Tier rangeBudget → Premium

🟠 Anthropic

Models9 (2 deprecated)
Cheapest input$1.00 (Haiku 4.5)
Most expensive output$75 (Claude 4 Opus)
Max context1M
Tier rangeBudget → Premium

🔵 Google

Models8 (2 deprecated)
Cheapest input$0.075 (Flash Lite)
Most expensive output$12 (Gemini 3.1 Pro)
Max context1M
Tier rangeBudget → Mid

🔴 DeepSeek

Models4 (1 deprecated)
Cheapest input$0.14 (V4 Flash)
Most expensive output$1.10 (V3)
Max context1M
Tier rangeBudget only

🟣 Mistral

Models3
Cheapest input$0.10 (Small 4)
Most expensive output$7.50 (Medium 3.5)
Max context262K
Tier rangeBudget → Mid

🟡 Meta (Together.ai)

Models4
Cheapest input$0.10 (Llama 3.1 8B)
Most expensive output$1.04 (Llama 3.3 70B)
Max context1M
Tier rangeBudget → Mid

⚪ Cohere

Models3
Cheapest input$0.50 (Command R)
Most expensive output$10 (Command A/R+)
Max context128K
Tier rangeBudget → Mid

🌙 xAI

Models2
Cheapest input$1.00 (Grok Build)
Most expensive output$2.50 (Grok 4.3)
Max context1M
Tier rangeBudget → Mid

4. Premium Models — What You're Paying For

The most expensive models command 2,400× the price of the cheapest. Here's what you get for the premium.

Model Provider Output $/1M vs Cheapest Context Best For
GPT-5.5 Pro OpenAI $180.00 1,800× 1.05M Research, complex reasoning
GPT-5.4 Pro OpenAI $180.00 1,800× 400K Enterprise, coding
Claude 4 Opus Anthropic $75.00 750× 200K Deprecated
Claude Fable 5 Anthropic $50.00 500× 1M Creative writing, storytelling
GPT-5.5 OpenAI $30.00 300× 1.05M General purpose, long context
Claude Opus 4.8 Anthropic $25.00 250× 1M Complex reasoning, analysis

💡 Is the premium worth it?

For most applications, no. Budget models (DeepSeek V4 Flash, Gemini Flash, Llama 4 Scout) handle 80-90% of use cases at 1/100th the cost. Reserve premium models for tasks that truly require advanced reasoning: multi-step planning, complex code generation, or nuanced analysis. Use a tiered routing strategy: simple requests → budget, complex → mid-tier, critical → premium.

5. Key Insights & Trends — July 2026

📈 Budget models are getting shockingly cheap

Sub-$0.10/1M input pricing is now standard for budget models. Google leads with Gemini 2.0 Flash Lite at $0.075/1M — that's 7.5 cents per million input tokens. At this price, processing 1 billion tokens costs just $75. This makes AI API costs essentially negligible for most applications.

🔄 The deprecation wave continues

Claude 4 Opus, Claude Sonnet 4, DeepSeek V3, Gemini 2.0 Flash, Gemini 2.0 Flash Lite, and Jamba 1.5 Large are all deprecated. If you're still using these models, migrate now — pricing and availability are not guaranteed. Check our Model Deprecation Checker for migration paths.

🌐 Context windows hit 1M tokens

Seven providers now offer 1M+ context windows: OpenAI (1.05M on GPT-5.5), Anthropic (1M on Opus/Sonnet/Haiku), Google (1M on all Gemini), DeepSeek (1M on V4), xAI (1M on Grok 4.3), and Together/Meta (1M on Llama 4). Long-context processing is now table stakes, not a premium feature.

💰 Output tokens are the real cost

Output tokens cost 3-6× more than input tokens across all providers. For a typical 3:1 input:output ratio, output tokens account for ~43% of your bill despite being only 25% of your tokens. Optimize by keeping responses short, using structured output formats, and avoiding verbose models for simple tasks.

6. Free Pricing API

All this data is available as a free, no-auth JSON API. Use it in your own tools, dashboards, or applications.

🔗 Endpoint

GET https://www.getapipulse.com/data/pricing.json

No authentication required. 88 models, 10 providers. Updated regularly. Full API docs →

Stop overpaying for AI APIs

Use our free calculator to see exactly how much you could save by switching models. Or run a full audit to get a personalized recommendation with migration code.

Related Tools