LLM API Cost Calculator
Pick a workload or enter your own numbers — you get the monthly bill on every model in the catalog, cheapest first.
No signup, no email
Workload
10,000 conversations a month, a long knowledge-base system prompt, short answers.
A fixed system prompt repeated on every request is billed cheaper by most providers. When a provider publishes no cache price, we bill that share at the full input rate — overstating the bill is safer than understating it.
What it will cost
The cheapest model for the “Customer support chatbot” workload is Mistral Small 3.2.2506 at $1.26 per month.
| Model | Per request | Per month | Saved by caching |
|---|---|---|---|
| Mistral Small 3.2.2506 | $0.00013 | $1.26 | — |
| Gemini 2.0 Flash Lite | $0.00014 | $1.40 | $0.41 |
| Gemini 2.0 Flash Lite 001 | $0.00014 | $1.40 | $0.41 |
| GPT-5 Nano | $0.00015 | $1.48 | $0.32 |
| Ministral 3 3b 2512 | $0.00015 | $1.50 | — |
| DeepSeek V4 Flash | $0.00015 | $1.53 | $0.99 |
| Gemini 2.5 Flash Lite | $0.00018 | $1.75 | $0.65 |
| Gemini Flash Lite | $0.00018 | $1.75 | $0.65 |
| Gemini 2.5 Flash Lite Preview 09.2025 | $0.00018 | $1.75 | $0.65 |
| Gemini 2.0 Flash 001 | $0.00019 | $1.86 | $0.54 |
| Gemini 2.5 Flash Lite Preview 06.17 | $0.00019 | $1.86 | $0.54 |
| Gemini 2.0 Flash | $0.00019 | $1.86 | $0.54 |
| GPT-4.1 Nano | $0.00019 | $1.86 | $0.54 |
| Devstral Small 2507 | $0.00021 | $2.10 | — |
| Devstral Small 2505 | $0.00021 | $2.10 | — |
179 × model
How this is calculated
Models with no published price (previews and non-text models) are excluded from the comparison: a zero in the source means “no price published”, not “free”. Cost = (input tokens ÷ 1,000,000) × input price + (output tokens ÷ 1,000,000) × output price, multiplied by request volume. The cached share of the input is billed at the cache-read price when the provider publishes one. Taxes, volume discounts and retries are not included — your real invoice will be slightly higher.
Prices come from our model catalog, synced daily with the open LiteLLM price reference.
Pricing verified August 15, 2026
Now the harder question
Counting tokens is the easy part. The hard part is deciding which processes in your company should go to AI agents, and in what order. That is what our AI-role org chart is for.
FAQ
- How many tokens is my text?
- A rough rule for English: one word ≈ 1.3 tokens, one page of text ≈ 500–700 tokens. Languages written in Cyrillic cost roughly two to three times more tokens for the same text. Only the model's own tokenizer gives an exact number.
- Why does the same model cost different amounts at different providers?
- We list the price charged by the company that made the model. Cloud resellers (Azure, Bedrock, Vertex) publish their own rates, which can differ by tens of percent in either direction.
- Are batch discounts included?
- No. Many providers charge about half price for asynchronous batch processing. If your workload tolerates delay, roughly halve the totals above.
en