RAG pipeline cost
Three separate bills in one system: index once (embedding your corpus), answer forever (LLM + retrieved context on every query), and keep it fresh (re-embedding churn + the vector DB).
The formula
corpus_tok = pages × words × 1.33
initial = corpus_tok/1M × E_in (one-time)
refresh = initial × refresh% (monthly)
per_query_in = chunks × chunkTok + query_tok
llm = queries × (per_query_in/1M × P_in + ans/1M × P_out) × 30
monthly = llm + refresh + vectorDB
Token prices as of . Re-ranking adds a second small model call per query — not modeled by default. Chunk overlap inflates indexed tokens ~10–15%; the refresh% input is where you absorb that.