Method & assumptions
Everything the calculators assume, in one place. Correct us and we'll fix it.
Where model prices come from
Per-token model prices are fetched daily, automatically from OpenRouter's public pricing API (/api/v1/models), converted to USD per 1M tokens, and rendered into calculators unchanged. The raw snapshot is diffed against the previous day; moves appear on the changes page. A second public source (LiteLLM's community price list) is fetched as a cross-check; divergences greater than 1% are flagged internally. If a day's cross-check fails, the table is still published with the primary source and date stamped.
Embedding models are a different case: the main model API does not list embedding endpoints at all, so the embedding prices in the price table and in the RAG / embedding calculators come from the second source (LiteLLM). If that fetch fails, the previous day's embedding list is reused unchanged and the date stamp shows it; if no list is available at all, those two calculators fall back to a manual $/1M rate input rather than silently assuming zero.
Which models appear in the calculators
Two deliberate exclusions, both visible as filters on the pricing table:
- Models that return images, audio or video are listed in the price table but kept out of every calculator dropdown. They bill on a different axis (per image, per audio second) and comparing them per token produces numbers that look plausible and are wrong.
- Batch and priority tiers (:batch variants and similar) are excluded, consistent with the "batch discounts deliberately excluded" rule below. Selecting one would quietly produce a discounted estimate while the page claims a conservative one.
Non-token rates (editable inputs)
Rates that aren't per-token are user-editable defaults. Each is marked with its source family and "as of" date on the calculator page. We deliberately do not hardcode vendor prices we can't re-verify daily — you should paste the number from your actual vendor invoice.
| Rate | Default | Basis |
|---|---|---|
| Speech-to-text | $0.005 / min | Range of mainstream hosted STT APIs (Whisper-class to Nova-class), 2026 |
| Text-to-speech | $0.015 / min | Mid-tier neural TTS, ~450 chars/min speech |
| Telephony (inbound) | $0.015 / min | Typical US SIP/voice-API inbound rate, excl. number rental |
| Email verification | $0.01 / lead | Volume tier of verification APIs |
| Vector database | $50 / month | Small-to-mid managed index; edit to your vendor |
| Words → tokens | ×1.33 | ≈0.75 words per token for English (BPE averages) |
What we deliberately exclude
- Prompt caching discounts — real but workload-specific; results without caching are the conservative case.
- Batch API discounts — same reason.
- Self-hosted GPU math — different problem class; these calculators are for hosted APIs.
The one number that matters
Industry breakdowns of production AI systems consistently find that LLM inference accounts for roughly 15–35% of total operating cost — the rest is speech, data infrastructure, retries, observability, and human review. That's why the Total Cost calculator exists: it shows the split, not just the token bill.