✨ All tools are free — no signup required. 88 models, 10 providers, full comparison.
GPT-5.6 Sol/Terra/Luna now live, Claude Sonnet 5 intro pricing ($2/$10 through Aug 31), and the best savings across 88 models from 10 providers. All tools free.
* GPT-oss is open-source (hosted inference, not OpenAI API). Cheapest API model: Gemini 2.5 Flash-Lite $0.10/1M.
The GPT-5.6 family is now live — Sol ($5/$30), Terra ($2.50/$15), and Luna ($1/$6), all with 1.05M context. Sol is the new flagship, matching GPT-5.5 pricing with larger context. Claude Sonnet 5 intro pricing ($2/$10) runs through Aug 31 — regular price jumps to $3/$15 after. The GPT-5.4 family remains the best value mid-tier (nano at $0.20/$1.25). GPT-5 and GPT-5 mini are deprecated (shutdown Dec 11 2026) — migrate now. Most developers are still overpaying by 60-90%.
| Model | Provider | Tier | Input / 1M | Output / 1M | Context |
|---|---|---|---|---|---|
| GPT-oss 20B | OpenAI | Budget | $0.08 | $0.35 | 128K |
| Gemini 2.5 Flash-Lite | Budget | $0.10 | $0.40 | 1M | |
| DeepSeek V4 Flash | DeepSeek | Budget | $0.14 | $0.28 | 1M |
| Mistral Small 4 | Mistral | Budget | $0.15 | $0.60 | 128K |
| Jamba Mini | AI21 | Budget | $0.20 | $0.40 | 256K |
Switch from expensive models to budget alternatives and save up to 98%.
| Switch From | Switch To | Output Savings |
|---|---|---|
|
GPT-5.5 Pro $180/1M output |
DeepSeek V4 Flash $0.28/1M output |
Save 99.8% |
|
GPT-5.5 $30/1M output |
DeepSeek V4 Flash $0.28/1M output |
Save 99.1% |
|
Claude Opus 4.8 $25/1M output |
Gemini 2.5 Flash-Lite $0.40/1M output |
Save 98.4% |
|
Claude Sonnet 5 $10/1M output intro |
Mistral Large 3 $1.50/1M output |
Save 85% |
|
GPT-5 $10/1M output |
DeepSeek V4 Pro $0.87/1M output |
Save 91.3% |
| Model | Status | Details | Replacement |
|---|---|---|---|
| GPT-5.6 Sol | ✨ New | $5.00/$30.00 per 1M tokens. 1.05M context. New OpenAI flagship, same price as GPT-5.5. | — |
| GPT-5.6 Terra | ✨ New | $2.50/$15.00 per 1M tokens. 1.05M context. Mid-tier workhorse. | — |
| GPT-5.6 Luna | ✨ New | $1.00/$6.00 per 1M tokens. 1.05M context. Budget-friendly with flagship context window. | — |
| Claude Sonnet 5 | ✨ Intro Price | $2/$10 per 1M tokens through Aug 31 (regular $3/$15). 1M context. Best mid-tier value right now. | — |
| Mistral Small 4 | ✨ New | $0.15/$0.60 per 1M tokens. 128K context. Ultra-budget option for high-volume tasks. | — |
| Grok Build 0.1 | 💰 Repriced | xAI code API pricing updated. Now $1.00/$2.00 per 1M tokens. 256K context. | — |
| GPT-4.1 | ✨ New | $2.00/$8.00 per 1M tokens. 1M context. Replaces GPT-4o at 20% cheaper pricing. | — |
| GPT-4.1 mini | ✨ New | $0.40/$1.60 per 1M tokens. 1M context. Budget-friendly replacement for GPT-4o mini. | — |
| GPT-5 | ⚠ Deprecated | Shutdown Dec 11, 2026. $1.25/$10.00. Migrate to GPT-5.5 before deadline. | GPT-5.5 |
| GPT-5 mini | ⚠ Deprecated | Shutdown Dec 11, 2026. $0.25/$2.00. Migrate to GPT-5.4 mini before deadline. | GPT-5.4 mini |
| GPT-4.1 nano | ⚠ Deprecated | Shutdown Oct 23, 2026. $0.10/$0.40. Migrate to GPT-5.4 nano before deadline. | GPT-5.4 nano |
| Llama 4 Scout / Maverick | 🚫 Delisted | Removed from Together.ai serverless Jul 2026. May still be available via dedicated endpoints. | Llama 3.3 70B |
| Jamba 1.5 Large | ⚠ Deprecated | Superseded by Jamba 1.7 Large with same pricing. | Jamba 1.7 Large |
| Claude Sonnet 4.6 | ⚠ Deprecated | Deprecated Jun 30, 2026. Same pricing as successor. | Claude Sonnet 5 |
| Claude 4 Opus | ⚠ Deprecated | Deprecated Jun 15, 2026. Was $15/$75 — 3× more expensive than successor. | Claude Opus 4.8 |
| Claude Sonnet 4 | ⚠ Deprecated | Deprecated Jun 15, 2026. Replaced by Sonnet 4.6, then Sonnet 5. | Claude Sonnet 5 |
| DeepSeek V3.2 | ⚠ Deprecated | Removed from DeepSeek pricing page Jul 7. No longer available. | DeepSeek V4 Flash |
| DeepSeek V3 | ⚠ Deprecated | Superseded by V4 Flash at lower pricing. | DeepSeek V4 Flash |
| Llama 3.1 70B | ⚠ Deprecated | Delisted from Together.ai serverless Jul 1. $0.88/$0.88. | Llama 3.3 70B |
| Llama 3.1 8B | ⚠ Deprecated | Delisted from Together.ai serverless Jul 1. $0.10/$0.10. | Llama 3.3 70B |
| Gemini 2.0 Flash / Flash Lite | ⚠ Deprecated | Replaced by Gemini 3 Flash and 3.1 Flash-Lite. | Gemini 3 Flash |
| GPT-4o / GPT-4o mini | ⚠ Deprecated | Deprecated Apr 2025. GPT-4.1 family is 20% cheaper with 8× larger context (1M vs 128K). | GPT-4.1 / nano |
Cheapest model per provider, sorted by output cost.
| Provider | Cheapest Model | Input / 1M | Output / 1M | Context |
|---|---|---|---|---|
| OpenAI | GPT-oss 20B | $0.08 | $0.35 | 128K |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | |
| DeepSeek | DeepSeek V4 Flash | $0.14 | $0.28 | 1M |
| Mistral | Mistral Small 4 | $0.15 | $0.60 | 128K |
| Cohere | Command R | $0.50 | $1.50 | 128K |
| xAI | Grok Build 0.1 | $1.00 | $2.00 | 256K |
| Meta | Llama 3.3 70B | $1.04 | $1.04 | 128K |
| AI21 | Jamba Mini | $0.20 | $0.40 | 256K |
| Moonshot | Kimi K2.6 | $0.95 | $4.00 | 256K |
| Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | 200K |
Use premium models (GPT-5.5, Opus 4.8) for complex reasoning and code generation. Use budget models (Flash, DeepSeek, Mistral Small) for classification, extraction, and simple Q&A. Splitting workloads across tiers typically saves 60-80% with no quality loss on routine tasks.
Get APIpulse for the full 67-model comparison, migration code snippets, PDF reports, and lifetime pricing updates. All tools are free.