AI API Cost per Request: Quick Reference Table

How much does a single API call actually cost? We calculated it for all 88 models across 10 providers at four common request sizes. Now with batch API pricing. Bookmark this page.

Assumption: Each request sends 3x more input tokens than output tokens (typical for chat, RAG, and code assistant workloads). Costs are per single request. All prices verified Jul 9, 2026. Batch mode shows 50% discounted rates for providers with batch APIs (OpenAI, Anthropic, Google) โ€” batch jobs process within 24 hours. Streaming mode adds 15% overhead to account for SSE framing and repeated context tokens in streamed responses.

All 59 Models โ€” Cost per Request

Sorted cheapest to most expensive. At 1K tokens, costs range from $0.000100 (Llama 3.1 8B) to $0.067500 (GPT-5.5 Pro) โ€” a 675x gap.

API Mode: Delivery:
Model Tier Provider 100 tok 500 tok 1K tok 5K tok
Llama 3.1 8BBudgetMeta (Together.ai)$0.000010$0.000050$0.000100$0.000500
GPT-oss 20BBudgetOpenAI$0.000015$0.000074$0.000148$0.000737
Llama 4 ScoutBudgetMeta (Together.ai)$0.000017$0.000084$0.000168$0.000838
Gemini 2.5 Flash-LiteBudgetGoogle$0.000017$0.000087$0.000175$0.000875
DeepSeek V4 FlashBudgetDeepSeek$0.000018$0.000087$0.000175$0.000875
Mistral Small 4BudgetMistral$0.000026$0.000131$0.000262$0.001313
GPT-4o miniBudgetOpenAI$0.000026$0.000131$0.000262$0.001313
GPT-oss 120BBudgetOpenAI$0.000026$0.000131$0.000262$0.001313
Llama 4 MaverickBudgetMeta (Together.ai)$0.000030$0.000150$0.000300$0.001500
DeepSeek V4 FlashBudgetDeepSeek$0.000048$0.000239$0.000478$0.002387
DeepSeek V4 ProBudgetDeepSeek$0.000055$0.000274$0.000548$0.002737
GPT-5 miniBudgetOpenAI$0.000070$0.000350$0.000700$0.003500
Command RBudgetCohere$0.000075$0.000375$0.000750$0.003750
Mistral Large 3BudgetMistral$0.000075$0.000375$0.000750$0.003750
Llama 3.3 70BMidMeta (Together.ai)$0.000104$0.000520$0.001040$0.005200
Claude Haiku 4.5BudgetAnthropic$0.000160$0.000800$0.001600$0.008000
Kimi K2.6BudgetMoonshot$0.000161$0.000806$0.001613$0.008063
Gemini 2.5 ProMidGoogle$0.000344$0.001719$0.003438$0.017188
Grok Build 0.1MidxAI$0.000350$0.001750$0.003500$0.017500
Jamba 1.5 LargeMidAI21$0.000350$0.001750$0.003500$0.017500
Command R+MidCohere$0.000438$0.002188$0.004375$0.021875
GPT-4oMidOpenAI$0.000438$0.002188$0.004375$0.021875
Gemini 3.1 ProMidGoogle$0.000450$0.002250$0.004500$0.022500
GPT-5.3 CodexMidOpenAI$0.000481$0.002406$0.004812$0.024063
Claude Sonnet 4.6MidAnthropic$0.000600$0.003000$0.006000$0.030000
Claude Sonnet 4.6MidAnthropic$0.000600$0.003000$0.006000$0.030000
Claude Opus 4.7PremiumAnthropic$0.001000$0.005000$0.010000$0.050000
GPT-5.5PremiumOpenAI$0.001125$0.005625$0.011250$0.056250
GPT-5PremiumOpenAI$0.001500$0.007500$0.015000$0.075000
Claude 4 OpusPremiumAnthropic$0.003000$0.015000$0.030000$0.150000
Grok 4.3PremiumxAI$0.006000$0.030000$0.060000$0.300000
GPT-5.5 ProPremiumOpenAI$0.006750$0.033750$0.067500$0.337500

Key Takeaways

The 675x Gap

The cheapest model (Llama 3.1 8B at $0.000100/request) costs 675x less than the most expensive (GPT-5.5 Pro at $0.067500/request) for a 1K-token request. At 5K tokens, the gap holds at 675x.

Calculate your exact monthly costs across all 88 models

Open the Calculator โ€” Free

How to Use This Table

These costs assume a 3:1 input-to-output token ratio. Your actual costs depend on your specific workload:

For exact calculations with your token ratios, use our interactive calculator or token estimator.

This was a snapshot. What about next month?
Prices change. New models launch. Our tools catch what a one-time calculation can't โ€” and saves you money every month.
Free Tools โ†’ ๐Ÿ” Free audit first

All Tools Are Free

No signup required to 67-model comparison, migration code snippets, PDF reports, price alerts, and cost monitoring. โœ… All tools free.

Free Tools โ†’
Free Tools โ€” Optimize Your Costs

Get model routing, caching strategies, and save 40%+ on API costs. 100% free โ€” no signup required.

Run Free Cost Audit โ†’ Free Tools โ†’

No signup required ยท 100% free