Skip to content

LLM API Cost Calculator

Pick a workload or enter your own numbers — you get the monthly bill on every model in the catalog, cheapest first.

No signup, no email

Workload

10,000 conversations a month, a long knowledge-base system prompt, short answers.

A fixed system prompt repeated on every request is billed cheaper by most providers. When a provider publishes no cache price, we bill that share at the full input rate — overstating the bill is safer than understating it.

What it will cost

The cheapest model for the “Customer support chatbot” workload is Ministral 3b at $0.85 per month.

ModelPer requestPer monthSaved by caching
Ministral 3b$0.000085$0.85$0.65
Ministral 3 3b 2512$0.000085$0.85$0.65
Ministral 3b 2512$0.000085$0.85$0.65
Mistral Small 3.2.2506$0.00013$1.26—
Ministral 8b 2512$0.00013$1.28$0.97
Ministral 8b$0.00013$1.28$0.97
Ministral 3 8b 2512$0.00013$1.28$0.97
Gemini 2.0 Flash Lite 001$0.00014$1.40$0.41
Gemini 2.0 Flash Lite$0.00014$1.40$0.41
Mistral Small$0.00015$1.45$0.65
Devstral Small$0.00015$1.45$0.65
GPT-5 Nano$0.00015$1.48$0.32
DeepSeek Coder$0.00016$1.61$0.91
Ministral 3 14b 2512$0.00017$1.70$1.30
Ministral 14b 2512$0.00017$1.70$1.30

224 × model

How this is calculated

Models with no published price (previews and non-text models) are excluded from the comparison: a zero in the source means “no price published”, not “free”. Cost = (input tokens ÷ 1,000,000) × input price + (output tokens ÷ 1,000,000) × output price, multiplied by request volume. The cached share of the input is billed at the cache-read price when the provider publishes one. Taxes, volume discounts and retries are not included — your real invoice will be slightly higher.

Prices come from our model catalog, synced daily with the open LiteLLM price reference.

Pricing verified September 29, 2026

Now the harder question

Counting tokens is the easy part. The hard part is deciding which processes in your company should go to AI agents, and in what order. That is what our AI-role org chart is for.

FAQ

How many tokens is my text?
A rough rule for English: one word ≈ 1.3 tokens, one page of text ≈ 500–700 tokens. Languages written in Cyrillic cost roughly two to three times more tokens for the same text. Only the model's own tokenizer gives an exact number.
Why does the same model cost different amounts at different providers?
We list the price charged by the company that made the model. Cloud resellers (Azure, Bedrock, Vertex) publish their own rates, which can differ by tens of percent in either direction.
Are batch discounts included?
No. Many providers charge about half price for asynchronous batch processing. If your workload tolerates delay, roughly halve the totals above.

en