Best value: OpenAI's text-embedding-3-small at $0.02/1M tokens is the cheapest option with excellent quality. For most RAG use cases, the small model is indistinguishable from the large model.

Real embedding costs

Let's say you have 10,000 documents averaging 2,000 tokens each (20M tokens total):

Embedding 10K documents (one-time)
OpenAI text-embedding-3-small$0.40
Google text-embedding-004$2.00
Cohere embed-v4$2.00

At $0.40 for 10K documents, embedding is practically free. The ongoing cost is embedding each query (~500 tokens), which costs fractions of a cent.

Component 2: Vector Storage & Search

This is where free tiers really shine. You have three options:

Option A: Fully Local (Free)

$0/mo

Use ChromaDB, FAISS, or SQLite-VSS on your own machine. Great for development and small datasets (under 100K documents). No monthly cost, but you manage infrastructure.

Option B: Free Cloud Tiers

$0/mo

Several vector databases offer generous free tiers:

Option C: Paid Vector DB

$25-70/mo

For production workloads beyond free tier limits:

For the cheapest setup, use ChromaDB locally during development and Pinecone's free tier in production. That keeps vector costs at $0.

Component 3: Generation Costs

This is the ongoing cost that grows with usage. Here's where model selection matters most. Let's compare costs for a typical RAG query: ~2,000 input tokens (retrieved context + query) and ~500 output tokens (answer).

Cost per RAG query (2K input + 500 output tokens)
Gemini 2.5 Flash-Lite$0.00030
DeepSeek V4 Flash$0.00042
Gemini 2.5 Flash-Lite$0.00045
GPT-4o mini$0.00060
DeepSeek V4 Pro$0.00132
GPT-5 mini$0.00150
Mistral Small 4$0.00060
Claude Haiku 4.5$0.00450
Claude Sonnet 4.6$0.01350
GPT-5$0.00750

The difference is dramatic. Gemini 2.5 Flash-Lite costs 45x less per query than Claude Sonnet 4.6. For RAG workloads where you're processing thousands of queries per day, this adds up fast.

Three Budget Tiers

Tier 1: Bootstrap ($0-5/mo)

Under $5/month

Best for: side projects, MVPs, internal tools

Total monthly cost at 1K queries/day: ~$1.50

Tier 2: Growth ($5-25/mo)

$5-25/month

Best for: production apps, startups with users

Total monthly cost at 5K queries/day: ~$20

Tier 3: Scale ($25-100/mo)

$25-100/month

Best for: SaaS products, high-traffic applications

Total monthly cost at 10K queries/day: ~$45-75

The Complete Cheapest RAG Stack

If your only goal is minimizing cost, here's the absolute cheapest production-ready RAG setup:

Monthly cost at 1,000 queries/day
Embedding (OpenAI small)$0.30/mo
Vector DB (ChromaDB local)$0.00/mo
Generation (Gemini 2.5 Flash-Lite)$1.35/mo
Total$1.65/mo

That's $1.65/month for a production RAG pipeline processing 1,000 queries per day. A year ago, the same setup would have cost $15-30/month.

Quality vs. Cost Tradeoffs

The cheapest models aren't always the best for RAG. Here's the quality spectrum:

Recommended: Hybrid routing

Use a cheap model (DeepSeek V4 Flash) for simple queries and route complex queries to a better model (GPT-5 mini or Claude Haiku). This cuts costs by 60-70% while maintaining quality for the queries that matter most.

Optimization Tips

  1. Chunk smartly โ€” Smaller chunks (256-512 tokens) mean less context sent to the LLM, reducing generation costs. Use overlapping chunks to maintain context.
  2. Cache common queries โ€” If 20% of your queries are repetitive, caching eliminates those generation costs entirely.
  3. Use metadata filtering โ€” Filter by date, category, or source before vector search to reduce the number of vectors compared.
  4. Batch embedding โ€” Embed documents in batches of 100+ for better throughput and lower per-token costs.
  5. Set max tokens โ€” Cap generation at the length you actually need. Shorter answers = lower costs.

Calculate your exact RAG costs โ€” Enter your document count, query volume, and token usage to see what you'd pay across every provider.

Calculate Your RAG Costs โ†’

โ€” See if you're overpaying for AI APIs

๐ŸŽฏ API Cost Score

Rate your API setup โ€” get a letter grade in 30 seconds

Related Reading

๐ŸŽฏ Rate Your API Setup in 30 Seconds

Get an A+ to F grade on your AI API costs. See how you compare and find cheaper alternatives instantly.

Get Your Cost Score โ†’

๐Ÿ“Š Generate Your Personalized API Cost Report

Select your model, enter your monthly spend, and get a custom savings report with cheaper alternatives โ€” free, in 60 seconds.

Found this useful? Share it with your team.

Want to optimize your AI API costs?

APIpulse includes free cost comparisons, exports, and recommendations that can save you up to 40%.

Free Tools โ†’

Save money: ๐Ÿ“Š Live API Pricing ยท Cost Optimizer โ€” find out how much you could save by switching models. Free tool.

๐Ÿ’ธ Looking for DeepSeek V4 Flash Alternatives?
5 models ranked by cost โ€” some offer better quality at similar prices.
See 5 DeepSeek V4 Flash Alternatives โ†’
๐Ÿ’ธ Looking for Sonnet 4.6 Alternatives?
5 models ranked by cost โ€” some are 90% cheaper.
See 5 Sonnet 4.6 Alternatives โ†’
๐Ÿ’ธ Looking for Mistral Small 4 Alternatives?
5 models ranked by cost โ€” some are 90% cheaper.
See 5 Mistral Small 4 Alternatives โ†’
๐Ÿ”ง Free Embeddable Pricing Widget
Add live AI API pricing to your docs, blog, or README with one script tag. 88 models, auto-updating.
Get the Free Widget โ†’ Free MCP Server โ†’