AI CostBase

RAG pipeline cost

Three separate bills in one system: index once (embedding your corpus), answer forever (LLM + retrieved context on every query), and keep it fresh (re-embedding churn + the vector DB).

The formula

corpus_tok = pages × words × 1.33 initial = corpus_tok/1M × E_in (one-time) refresh = initial × refresh% (monthly) per_query_in = chunks × chunkTok + query_tok llm = queries × (per_query_in/1M × P_in + ans/1M × P_out) × 30 monthly = llm + refresh + vectorDB

Token prices as of . Re-ranking adds a second small model call per query — not modeled by default. Chunk overlap inflates indexed tokens ~10–15%; the refresh% input is where you absorb that.