What changed in AI API pricing — and All tools are now free.
OpenAI launched GPT-5.4 mini at $0.75 input / $4.50 output per 1M tokens. Priced between GPT-4o mini and GPT-4o, it delivers significantly better reasoning than its predecessor while remaining OpenAI's most cost-effective new-gen model. For chatbots, content generation, and data extraction that need stronger reasoning than GPT-4o mini could offer.
The new flagship from OpenAI. GPT-5.4 Pro targets complex reasoning, code generation, and agentic workflows at $30.00 input / $180.00 output per 1M tokens. Premium pricing positions it alongside Claude Opus 4.8 and GPT-5.5 Pro for the most demanding workloads.
Google's answer to the price war. Gemini 3.1 Flash-Lite at $0.25 input / $1.50 output per 1M tokens is one of the cheapest multimodal models from a major provider. Perfect for high-volume tasks like classification, routing, and simple Q&A where cost matters most.
DeepSeek's latest at $0.14 input / $0.28 output per 1M tokens. Remarkably cheap for a model with strong reasoning capabilities. If you're using GPT-4o for tasks that don't need its full capability, DeepSeek V4 Flash could cut your costs by 90%+.
OpenAI enters the budget open-weight space. GPT-oss 120B at $0.15 input / $0.60 output per 1M tokens. Same price as GPT-4o mini but with a different architecture trade-off. Worth benchmarking against your current model for classification and extraction tasks.
xAI rebranded Grok 3 → Grok 4.3 and slashed pricing from $3.00/$15.00 to $1.25/$2.50 per 1M tokens. That's an 83% output price cut. At $2.50/1M output, Grok 4.3 is now cheaper than Claude Haiku ($5) and competitive with Gemini 3 Flash ($3) for tasks that need solid reasoning. If you dismissed xAI's pricing before, it's time to take another look.
Major cleanup across providers. Anthropic: Claude 4 Opus → Opus 4.8, Sonnet 4.6 → Sonnet 5, Sonnet 4 → Sonnet 4.6. Google: Gemini 2.0 Flash → 3 Flash, Gemini 2.0 Flash Lite → 3.1 Flash-Lite. DeepSeek: V3 → V4 Flash. AI21: Jamba 1.5 → 1.7. If your code references any of these endpoints, you need to migrate. APIpulse's audit tool now warns about deprecated models and shows migration paths.
Twelve months ago, the cheapest API model from a major provider was ~$0.15/1M input tokens. Today, GPT-oss 20B is at $0.08, Mistral Small at $0.10, and Gemini 2.5 Flash-Lite at $0.10. The "good enough for most tasks" price has nearly halved in a year. If you locked in pricing assumptions 6 months ago, you're overpaying.
While budget models race to the bottom, premium-tier pricing (GPT-5.4 Pro, Claude Opus 4.8, GPT-5.5 Pro) remains at $5-30 input / $25-180 output per 1M tokens. The gap between "cheap" and "best" is now 30-600x. This creates a clear optimization opportunity: route simple tasks to cheap models, reserve premium for complex reasoning.
Run a free 30-second audit of your current API model. See exactly how much you could save — and get migration code if you need to switch.
Run My Free Audit →Get the API Pricing Digest delivered every Friday. No spam — just pricing intelligence.
Free. Unsubscribe anytime. We respect your inbox.
APIpulse monitors every model across 10 providers. Get alerts when prices drop, migration code ready to paste, and a cost dashboard to track savings. No signup required — everything is free.
Free Tools →