📊 AI API Pricing Cheat Sheet 2026

Every model. Every provider. One page. Updated weekly.

Last updated: July 7, 2026 · 88 models · 10 providers

✓ 38 Active Models 10 Providers 37× Price Range
$0.07
Cheapest Input (per 1M tok)
$0.10
Cheapest Output (per 1M tok)
$180
Most Expensive Output
1.05M
Largest Context Window

📋 Full Pricing Table — All 59 Models

Click any column header to sort. Prices are per 1 million tokens. Filter by provider or tier.

Model Provider Tier Input $/1M Output $/1M Blended* Context

* Blended = 3:1 output:input ratio (typical for chat). Click column headers to sort. Greyed rows are deprecated.

🏢 Provider Comparison

At a glance: what each provider offers and where they shine.

OpenAI

9 models (7 active)
Range: $0.08 — $180/1M output
Strength: Broadest lineup
GPT-oss 20B → GPT-5.5 Pro

Anthropic

6 models (4 active)
Range: $5 — $75/1M output
Strength: Code & reasoning
Haiku 4.5 → Claude 4 Opus

Google

8 models (6 active)
Range: $0.30 — $12/1M output
Strength: Best value mid-tier
Flash Lite → Gemini 3.1 Pro

DeepSeek

4 models (3 active)
Range: $0.28 — $0.87/1M output
Strength: Cheapest quality
V4 Flash → V4 Pro

Mistral

3 models (3 active)
Range: $0.30 — $7.50/1M output
Strength: EU compliance
Small 4 → Large 3

Meta (via Together)

4 models (3 active)
Range: $0.10 — $0.88/1M output
Strength: Open-source pricing
Llama 3.1 8B → Llama 4

Cohere

3 models (3 active)
Range: $1.50 — $10/1M output
Strength: Enterprise RAG
Command R → Command A

xAI

2 models (2 active)
Range: $0.50 — $2.50/1M output
Strength: Large context
Grok Build → Grok 4.3

Moonshot

1 model
Range: $4.00/1M output
Strength: Chinese market
Kimi K2.6

AI21

2 models (1 active)
Range: $8.00/1M output
Strength: Long-context SSM
Jamba 1.7 Large

💰 Cost Per Task — What You'll Actually Pay

Real-world cost estimates for common AI tasks. Based on typical token counts.

💬 Simple Chat

~500 input + ~300 output tokens
Llama 3.1 8B$0.0001
GPT-5 mini$0.001
Claude Sonnet 4.6$0.006
GPT-5$0.004
Claude Opus 4.8$0.010

📄 Summarize Document

~4,000 input + ~500 output tokens
Llama 3.1 8B$0.0005
GPT-5 mini$0.002
Claude Sonnet 4.6$0.020
GPT-5$0.010
Claude Opus 4.8$0.033

🔧 Generate Code

~1,000 input + ~2,000 output tokens
Llama 3.1 8B$0.0003
GPT-5 mini$0.004
Claude Sonnet 4.6$0.033
GPT-5$0.021
Claude Opus 4.8$0.055

📊 Analyze Data

~10,000 input + ~1,000 output tokens
Llama 3.1 8B$0.001
GPT-5 mini$0.005
Claude Sonnet 4.6$0.045
GPT-5$0.023
Claude Opus 4.8$0.075

🤖 AI Agent Step

~2,000 input + ~1,500 output tokens
Llama 3.1 8B$0.0004
GPT-5 mini$0.004
Claude Sonnet 4.6$0.029
GPT-5$0.018
Claude Opus 4.8$0.048

📚 RAG Query

~15,000 input + ~500 output tokens
Llama 3.1 8B$0.002
GPT-5 mini$0.005
Claude Sonnet 4.6$0.053
GPT-5$0.024
Claude Opus 4.8$0.088

💡 Key insight: Output tokens cost 3–20× more than input tokens across all providers. Design your prompts to minimize output length where possible.

🎯 Quick Recommendations by Use Case

Not sure which model to pick? Start here.

💸 Lowest Cost

When budget is everything
Llama 3.1 8B
$0.10 / $0.10 per 1M

⚖️ Best Value

Quality + affordability
DeepSeek V4 Pro
$0.43 / $0.87 per 1M

🧠 Best Reasoning

Complex analysis & coding
Claude Opus 4.8
$5.00 / $25.00 per 1M

⚡ Best Speed

Real-time apps, chatbots
Gemini 3.5 Flash
$1.50 / $9.00 per 1M

📝 Long Documents

Large context window needed
GPT-5.5
$5.00 / $30.00 per 1M (1.05M ctx)

💻 Code Generation

Coding copilots & IDE tools
Claude Sonnet 4.6
$3.00 / $15.00 per 1M

🏢 Enterprise RAG

Retrieval-augmented generation
Command R+
$2.50 / $10.00 per 1M

🌍 EU Compliance

GDPR-friendly, EU hosting
Mistral Large 3
$0.50 / $1.50 per 1M

🚀 Budget General

Good enough for most tasks
GPT-5 mini
$0.25 / $2.00 per 1M

⚠️ 3 Hidden Costs Most Developers Miss

The per-token price is just the beginning. Here's what catches teams off guard.

1. Output Token Markup (3–20×)

Every provider charges more for output tokens than input. The markup ranges from 3× (DeepSeek) to 20× (GPT-5.5 Pro). For code generation and long-form writing, output tokens dominate your bill.

💡 Tip: Set max_tokens limits. Use structured output (JSON mode) to reduce verbose responses.

2. Context Window Waste

Sending full conversation history on every API call means your input costs grow linearly with conversation length. A 50-turn chat with GPT-5 at 10K context costs $0.125 per message just for input.

💡 Tip: Implement conversation summarization. Trim old messages. Use sliding window approaches.

3. Retry & Error Costs

Rate limits, timeouts, and malformed responses all consume tokens. With aggressive retry logic, you can accidentally 3× your costs. Some providers charge for partial completions.

💡 Tip: Implement exponential backoff. Set retry budgets. Monitor error rates per provider.

Want detailed cost reports for your exact usage?

APIpulse includes monthly cost tracking, provider comparison reports, and optimization recommendations tailored to your API usage patterns.

Explore Free Tools →

100% free · No signup · No credit card