RAG System Cost Calculator
Estimate vector DB storage, embedding, and query costs for a Retrieval-Augmented Generation system.
About this calculator
This calculator walks through a RAG pipeline's cost stack in the order money actually gets spent. First, chunking: your document count and average tokens-per-document get split into overlapping chunks of the size you specify, using (tokens − overlap) / (chunk size − overlap) to find how many chunks each document yields — larger overlap means more chunks (and more redundant embedding cost) for the same document, since consecutive chunks re-embed a shared slice of text to preserve context across boundaries. Every one of those chunks gets a one-time embedding cost, calculated here at OpenAI's ada-002 rate of $0.10 per million tokens regardless of what embedding-dimensions value you enter — so if you set 3072 dimensions to model text-embedding-3-large, the storage math correctly grows (each vector costs dimensions × 4 bytes for float32 plus 50 bytes of metadata overhead) but the embedding-cost line still assumes ada-002 pricing, a real limitation to keep in mind when comparing providers.
Vector database hosting is approximated linearly at $0.07/month per 1,000 stored vectors, modeled on Pinecone's roughly $70/month per million vectors. Ongoing query cost has two parts: embedding each incoming query (same ada-002 rate) and an LLM synthesis call sized by Top-K retrieved chunks × chunk size as input tokens plus a fixed 500 output tokens, priced at GPT-4o's published $2.50/1M input and $10/1M output rates. Because the embedding and LLM pricing are both hard-coded to specific providers, use this as a directional budget for a given architecture, then substitute your actual provider's rates for a precise number.
Inputs
Results
Total chunks
50,000
Total monthly cost ($)
$347.04
How to Use This Calculator
- Enter Document Count and average Tokens per Document for your knowledge base (a page is ~500 tokens).
- Set Chunk Size (512 tokens is common) and Chunk Overlap to control how documents are split for retrieval.
- Enter Embedding Dimensions (1536 for OpenAI ada-002) and expected Queries per Day.
- Set Top-K Retrieved Chunks to control how many context chunks are injected per LLM call.
- Review Total Chunks, One-time Embedding Cost, Vector Storage (GB), Monthly Query Cost, and Total Monthly Cost to design your RAG system budget.
How the result changes with Number of documents
| Number of documents | Total chunks | Total monthly cost ($) |
|---|---|---|
| 5,000 | 25,000 | $345.29 |
| 7,500 | 37,500 | $346.16 |
| 15,000 | 75,000 | $348.79 |
| 25,000 | 125,000 | $352.29 |
What each input means
- Number of documents
- Total documents in your knowledge base corpus.
- Avg tokens per document
- Average token length per document. A typical page of text is ~500 tokens.
- Chunk size (tokens)
- Size of each text chunk for embedding. Common sizes: 256, 512, 1024.
- Chunk overlap (tokens)
- Token overlap between consecutive chunks to preserve context across boundaries.
- Embedding dimensions
- Vector dimensions. OpenAI ada-002: 1536, text-embedding-3-large: 3072, Cohere: 1024.
- Queries per day
- Expected number of RAG queries per day from your application.
- Top-K retrieved chunks
- Number of most-similar chunks retrieved per query for LLM context.
What each result means
- Total chunks
- Total text chunks generated from your document corpus.
- One-time embedding cost ($)
- Cost to embed all chunks using OpenAI ada-002 at $0.10/1M tokens.
- Vector storage (GB)
- Estimated vector database storage in gigabytes (float32 + metadata).
- Vector DB monthly cost ($)
- Estimated Pinecone-equivalent hosting cost (~$0.07/month per 1K vectors).
- Daily query cost ($)
- Daily cost for query embeddings + LLM synthesis (GPT-4o pricing).
- Monthly query cost ($)
- Projected 30-day query cost.
- Total monthly cost ($)
- Vector DB hosting + query costs combined.
- Embedding tokens (millions)
- Total tokens sent to the embedding API during indexing.
How this is calculated
Worked example, using the default values
- Identify Input Parameters4 parametersNumber of documents = 10000, Avg tokens per document = 2000, Chunk size (tokens) = 512, Chunk overlap (tokens) = 50 = 7 input(s) provided
- Calculate Total chunksTotal chunks = documentCount * chunksPerDoc50000 = 50000
- Calculate Total monthly costTotal monthly cost = vectorDbMonthlyCost + monthlyQueryCost347.04 = $347.04
- Calculate One-time embedding costOne-time embedding cost = (totalEmbeddingTokens * 0.10) / 1_000_0002.56 = $2.56
- Calculate Vector storageVector storage = totalStorageBytes / (1024 ^ 3)0.288 = 0.288
Engine last updated . Checked against 3 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.
Frequently Asked Questions
Why doesn't the embedding cost change when I raise Embedding Dimensions to model a different provider?
The embedding cost line is hard-coded to OpenAI ada-002's $0.10-per-million-token rate regardless of the Embedding Dimensions value you enter, so raising dimensions to 3072 to model text-embedding-3-large changes the Vector Storage figure (since storage scales with dimensions × 4 bytes per vector) but leaves embeddingCost untouched. Treat Embedding Dimensions as controlling storage size only, and substitute your actual provider's per-token rate if it differs from ada-002.
How does increasing Chunk Overlap affect Total Chunks and cost?
Chunks per document come from (tokens − overlap) / (chunk size − overlap), so a larger overlap shrinks the effective stride between chunks and produces more chunks per document — and since every chunk gets embedded and stored, more overlap directly increases embedding cost, storage, and vector DB hosting cost even though your underlying documents haven't changed.
What LLM is assumed for the query cost calculation, and can I adjust it?
Query synthesis cost is priced at GPT-4o's published rates: $2.50 per million input tokens and $10 per million output tokens, with output fixed at 500 tokens per query and input equal to Top-K × Chunk Size. There's no input to change the model, so if you're using a different LLM for synthesis, scale the Daily/Monthly Query Cost figures by your model's relative pricing.
How is Vector Storage (GB) calculated?
Each stored vector costs Embedding Dimensions × 4 bytes (for float32 precision) plus a fixed 50 bytes of metadata overhead per vector, multiplied by Total Chunks and converted from bytes to gigabytes. Because dimensions scale storage linearly, doubling Embedding Dimensions roughly doubles the storage estimate for the same document set.
Related Calculators
The questions that sit next to this one — chosen by subject, including calculators filed under a different category.
Embedding Cost Calculator
Calculate embedding API costs and storage requirements across OpenAI, Cohere, and self-hosted models.
Ai ToolsLLM Token Calculator
Estimate token count and API cost from text length across different tokenizers (GPT-4, Claude, Llama).
Ai ToolsAI Model Cost Comparison Calculator
Compare per-token and per-request costs across AI providers including GPT-4o, Claude, Gemini, and self-hosted Llama.
More in Technology & Computing.