✨ All tools are free — no signup required. 88 models, 10 providers, full comparison.

📊 Weekly Report

AI API Pricing Report

GPT-5.6 Sol/Terra/Luna now live, Claude Sonnet 5 intro pricing ($2/$10 through Aug 31), and the best savings across 88 models from 10 providers. All tools free.

📅 Week of July 14, 2026 🔄 Updated Jul 12, 2026 📈 88 models tracked
67
Models Tracked
10
Providers
$0.08
Cheapest Input *
98%
Max Savings

* GPT-oss is open-source (hosted inference, not OpenAI API). Cheapest API model: Gemini 2.5 Flash-Lite $0.10/1M.

💡 This Week's Key Insight

The GPT-5.6 family is now live — Sol ($5/$30), Terra ($2.50/$15), and Luna ($1/$6), all with 1.05M context. Sol is the new flagship, matching GPT-5.5 pricing with larger context. Claude Sonnet 5 intro pricing ($2/$10) runs through Aug 31 — regular price jumps to $3/$15 after. The GPT-5.4 family remains the best value mid-tier (nano at $0.20/$1.25). GPT-5 and GPT-5 mini are deprecated (shutdown Dec 11 2026) — migrate now. Most developers are still overpaying by 60-90%.

🏆 Best Value Models by Tier

Model Provider Tier Input / 1M Output / 1M Context
GPT-oss 20B OpenAI Budget $0.08 $0.35 128K
Gemini 2.5 Flash-Lite Google Budget $0.10 $0.40 1M
DeepSeek V4 Flash DeepSeek Budget $0.14 $0.28 1M
Mistral Small 4 Mistral Budget $0.15 $0.60 128K
Jamba Mini AI21 Budget $0.20 $0.40 256K

💰 Biggest Savings Opportunities

Switch from expensive models to budget alternatives and save up to 98%.

Switch From Switch To Output Savings
GPT-5.5 Pro
$180/1M output
DeepSeek V4 Flash
$0.28/1M output
Save 99.8%
GPT-5.5
$30/1M output
DeepSeek V4 Flash
$0.28/1M output
Save 99.1%
Claude Opus 4.8
$25/1M output
Gemini 2.5 Flash-Lite
$0.40/1M output
Save 98.4%
Claude Sonnet 5
$10/1M output intro
Mistral Large 3
$1.50/1M output
Save 85%
GPT-5
$10/1M output
DeepSeek V4 Pro
$0.87/1M output
Save 91.3%

Recent Model Changes

Model Status Details Replacement
GPT-5.6 Sol ✨ New $5.00/$30.00 per 1M tokens. 1.05M context. New OpenAI flagship, same price as GPT-5.5.
GPT-5.6 Terra ✨ New $2.50/$15.00 per 1M tokens. 1.05M context. Mid-tier workhorse.
GPT-5.6 Luna ✨ New $1.00/$6.00 per 1M tokens. 1.05M context. Budget-friendly with flagship context window.
Claude Sonnet 5 ✨ Intro Price $2/$10 per 1M tokens through Aug 31 (regular $3/$15). 1M context. Best mid-tier value right now.
Mistral Small 4 ✨ New $0.15/$0.60 per 1M tokens. 128K context. Ultra-budget option for high-volume tasks.
Grok Build 0.1 💰 Repriced xAI code API pricing updated. Now $1.00/$2.00 per 1M tokens. 256K context.
GPT-4.1 ✨ New $2.00/$8.00 per 1M tokens. 1M context. Replaces GPT-4o at 20% cheaper pricing.
GPT-4.1 mini ✨ New $0.40/$1.60 per 1M tokens. 1M context. Budget-friendly replacement for GPT-4o mini.
GPT-5 ⚠ Deprecated Shutdown Dec 11, 2026. $1.25/$10.00. Migrate to GPT-5.5 before deadline. GPT-5.5
GPT-5 mini ⚠ Deprecated Shutdown Dec 11, 2026. $0.25/$2.00. Migrate to GPT-5.4 mini before deadline. GPT-5.4 mini
GPT-4.1 nano ⚠ Deprecated Shutdown Oct 23, 2026. $0.10/$0.40. Migrate to GPT-5.4 nano before deadline. GPT-5.4 nano
Llama 4 Scout / Maverick 🚫 Delisted Removed from Together.ai serverless Jul 2026. May still be available via dedicated endpoints. Llama 3.3 70B
Jamba 1.5 Large ⚠ Deprecated Superseded by Jamba 1.7 Large with same pricing. Jamba 1.7 Large
Claude Sonnet 4.6 ⚠ Deprecated Deprecated Jun 30, 2026. Same pricing as successor. Claude Sonnet 5
Claude 4 Opus ⚠ Deprecated Deprecated Jun 15, 2026. Was $15/$75 — 3× more expensive than successor. Claude Opus 4.8
Claude Sonnet 4 ⚠ Deprecated Deprecated Jun 15, 2026. Replaced by Sonnet 4.6, then Sonnet 5. Claude Sonnet 5
DeepSeek V3.2 ⚠ Deprecated Removed from DeepSeek pricing page Jul 7. No longer available. DeepSeek V4 Flash
DeepSeek V3 ⚠ Deprecated Superseded by V4 Flash at lower pricing. DeepSeek V4 Flash
Llama 3.1 70B ⚠ Deprecated Delisted from Together.ai serverless Jul 1. $0.88/$0.88. Llama 3.3 70B
Llama 3.1 8B ⚠ Deprecated Delisted from Together.ai serverless Jul 1. $0.10/$0.10. Llama 3.3 70B
Gemini 2.0 Flash / Flash Lite ⚠ Deprecated Replaced by Gemini 3 Flash and 3.1 Flash-Lite. Gemini 3 Flash
GPT-4o / GPT-4o mini ⚠ Deprecated Deprecated Apr 2025. GPT-4.1 family is 20% cheaper with 8× larger context (1M vs 128K). GPT-4.1 / nano

📊 Provider Comparison — Budget Tier

Cheapest model per provider, sorted by output cost.

Provider Cheapest Model Input / 1M Output / 1M Context
OpenAI GPT-oss 20B $0.08 $0.35 128K
Google Gemini 2.5 Flash-Lite $0.10 $0.40 1M
DeepSeek DeepSeek V4 Flash $0.14 $0.28 1M
Mistral Mistral Small 4 $0.15 $0.60 128K
Cohere Command R $0.50 $1.50 128K
xAI Grok Build 0.1 $1.00 $2.00 256K
Meta Llama 3.3 70B $1.04 $1.04 128K
AI21 Jamba Mini $0.20 $0.40 256K
Moonshot Kimi K2.6 $0.95 $4.00 256K
Anthropic Claude Haiku 4.5 $1.00 $5.00 200K

🎯 Quick Tip: Right-Size Your Models

Use premium models (GPT-5.5, Opus 4.8) for complex reasoning and code generation. Use budget models (Flash, DeepSeek, Mistral Small) for classification, extraction, and simple Q&A. Splitting workloads across tiers typically saves 60-80% with no quality loss on routine tasks.

📊 Stop Guessing Your AI Costs

Get APIpulse for the full 67-model comparison, migration code snippets, PDF reports, and lifetime pricing updates. All tools are free.

📚 More Resources

📈 Pricing Trends
Historical price data & charts
⚠️ Deprecation Tracker
Which models are retiring
🧮 Cost Calculator
Estimate your monthly spend
🤖 AI Advisor
Find the best model for your use case
🔄 GPT-5.4 mini Alternatives
5 cheaper options for the new model
🔄 Gemini 3.1 Flash-Lite Alts
Compare ultra-budget models

Related Tools