Skip to main content
Calcimator

RAG System Cost Calculator

Estimate vector DB storage, embedding, and query costs for a Retrieval-Augmented Generation system.

About this calculator

This calculator walks through a RAG pipeline's cost stack in the order money actually gets spent. First, chunking: your document count and average tokens-per-document get split into overlapping chunks of the size you specify, using (tokens − overlap) / (chunk size − overlap) to find how many chunks each document yields — larger overlap means more chunks (and more redundant embedding cost) for the same document, since consecutive chunks re-embed a shared slice of text to preserve context across boundaries. Every one of those chunks gets a one-time embedding cost, calculated here at OpenAI's ada-002 rate of $0.10 per million tokens regardless of what embedding-dimensions value you enter — so if you set 3072 dimensions to model text-embedding-3-large, the storage math correctly grows (each vector costs dimensions × 4 bytes for float32 plus 50 bytes of metadata overhead) but the embedding-cost line still assumes ada-002 pricing, a real limitation to keep in mind when comparing providers.

Vector database hosting is approximated linearly at $0.07/month per 1,000 stored vectors, modeled on Pinecone's roughly $70/month per million vectors. Ongoing query cost has two parts: embedding each incoming query (same ada-002 rate) and an LLM synthesis call sized by Top-K retrieved chunks × chunk size as input tokens plus a fixed 500 output tokens, priced at GPT-4o's published $2.50/1M input and $10/1M output rates. Because the embedding and LLM pricing are both hard-coded to specific providers, use this as a directional budget for a given architecture, then substitute your actual provider's rates for a precise number.

Inputs

Results

Total chunks

50,000

Total monthly cost ($)

$347.04

One-time embedding cost ($)$2.56
Vector storage (GB)0.29
Vector DB monthly cost ($)$3.50
Daily query cost ($)$11.45
Monthly query cost ($)$343.54
Embedding tokens (millions)25.6
How to Use This Calculator
  1. Enter Document Count and average Tokens per Document for your knowledge base (a page is ~500 tokens).
  2. Set Chunk Size (512 tokens is common) and Chunk Overlap to control how documents are split for retrieval.
  3. Enter Embedding Dimensions (1536 for OpenAI ada-002) and expected Queries per Day.
  4. Set Top-K Retrieved Chunks to control how many context chunks are injected per LLM call.
  5. Review Total Chunks, One-time Embedding Cost, Vector Storage (GB), Monthly Query Cost, and Total Monthly Cost to design your RAG system budget.

How the result changes with Number of documents

Number of documentsTotal chunksTotal monthly cost ($)
5,00025,000$345.29
7,50037,500$346.16
15,00075,000$348.79
25,000125,000$352.29

What each input means

Number of documents
Total documents in your knowledge base corpus.
Avg tokens per document
Average token length per document. A typical page of text is ~500 tokens.
Chunk size (tokens)
Size of each text chunk for embedding. Common sizes: 256, 512, 1024.
Chunk overlap (tokens)
Token overlap between consecutive chunks to preserve context across boundaries.
Embedding dimensions
Vector dimensions. OpenAI ada-002: 1536, text-embedding-3-large: 3072, Cohere: 1024.
Queries per day
Expected number of RAG queries per day from your application.
Top-K retrieved chunks
Number of most-similar chunks retrieved per query for LLM context.

What each result means

Total chunks
Total text chunks generated from your document corpus.
One-time embedding cost ($)
Cost to embed all chunks using OpenAI ada-002 at $0.10/1M tokens.
Vector storage (GB)
Estimated vector database storage in gigabytes (float32 + metadata).
Vector DB monthly cost ($)
Estimated Pinecone-equivalent hosting cost (~$0.07/month per 1K vectors).
Daily query cost ($)
Daily cost for query embeddings + LLM synthesis (GPT-4o pricing).
Monthly query cost ($)
Projected 30-day query cost.
Total monthly cost ($)
Vector DB hosting + query costs combined.
Embedding tokens (millions)
Total tokens sent to the embedding API during indexing.

How this is calculated

Worked example, using the default values

  1. Identify Input Parameters
    4 parameters
    Number of documents = 10000, Avg tokens per document = 2000, Chunk size (tokens) = 512, Chunk overlap (tokens) = 50 = 7 input(s) provided
  2. Calculate Total chunks
    Total chunks = documentCount * chunksPerDoc
    50000 = 50000
  3. Calculate Total monthly cost
    Total monthly cost = vectorDbMonthlyCost + monthlyQueryCost
    347.04 = $347.04
  4. Calculate One-time embedding cost
    One-time embedding cost = (totalEmbeddingTokens * 0.10) / 1_000_000
    2.56 = $2.56
  5. Calculate Vector storage
    Vector storage = totalStorageBytes / (1024 ^ 3)
    0.288 = 0.288

Engine last updated . Checked against 3 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Why doesn't the embedding cost change when I raise Embedding Dimensions to model a different provider?

The embedding cost line is hard-coded to OpenAI ada-002's $0.10-per-million-token rate regardless of the Embedding Dimensions value you enter, so raising dimensions to 3072 to model text-embedding-3-large changes the Vector Storage figure (since storage scales with dimensions × 4 bytes per vector) but leaves embeddingCost untouched. Treat Embedding Dimensions as controlling storage size only, and substitute your actual provider's per-token rate if it differs from ada-002.

How does increasing Chunk Overlap affect Total Chunks and cost?

Chunks per document come from (tokens − overlap) / (chunk size − overlap), so a larger overlap shrinks the effective stride between chunks and produces more chunks per document — and since every chunk gets embedded and stored, more overlap directly increases embedding cost, storage, and vector DB hosting cost even though your underlying documents haven't changed.

What LLM is assumed for the query cost calculation, and can I adjust it?

Query synthesis cost is priced at GPT-4o's published rates: $2.50 per million input tokens and $10 per million output tokens, with output fixed at 500 tokens per query and input equal to Top-K × Chunk Size. There's no input to change the model, so if you're using a different LLM for synthesis, scale the Daily/Monthly Query Cost figures by your model's relative pricing.

How is Vector Storage (GB) calculated?

Each stored vector costs Embedding Dimensions × 4 bytes (for float32 precision) plus a fixed 50 bytes of metadata overhead per vector, multiplied by Total Chunks and converted from bytes to gigabytes. Because dimensions scale storage linearly, doubling Embedding Dimensions roughly doubles the storage estimate for the same document set.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.