AI & Machine Learning
AI Inference Cost Calculator
Turn token counts and per-million prices into cost per request, per month and per year for an LLM API.
Free to useNo sign-up requiredNo watermarkRuns in your browser
Last reviewed: 30 September 2026 by Vishal Senthilkumar
LLM APIs charge per token, with separate prices for the tokens you send and the tokens the model writes - and output usually costs several times more than input. That makes it hard to see, from a price table alone, what a feature will cost once real users are on it.
Enter your traffic and the typical size of a request and response, and this calculator gives the cost per request, per day, per month and per year, the cost per 1,000 requests, and a single blended price per million tokens for your particular input-output mix.
How this tool works
Enter your traffic
Requests per day and how many days a month the feature runs.
Enter typical token counts
Input tokens per request (including the system prompt and context) and output tokens per response.
Enter prices
Input and output prices per 1M tokens, or pick a preset.
Read the costs
Per request, per day, per month, per year, per 1,000 requests and blended per 1M tokens.
How it works
The cost of one request is input tokens × input price ÷ 1,000,000 plus output tokens × output price ÷ 1,000,000. Daily cost multiplies that by requests per day, monthly by days per month, and annual cost is twelve months.
The blended cost per 1M tokens divides the cost of a request by all the tokens in it. Because output is priced higher, a workload with long answers has a blended price close to the output price; a workload with long prompts and short answers sits close to the input price.
The share of cost from output tells you which lever matters more. If output dominates, shorter answers and a sensible maximum output length help most; if input dominates, trimming context and prompt caching do.
Common use cases
- Budgeting a new chatbot, assistant or AI feature before launch.
- Comparing two models by entering each one’s prices against the same traffic.
- Setting a price or usage limit for an AI feature in your own product.
- Checking whether a cheaper model or shorter prompts are worth the effort.
Getting realistic token counts
Input is more than the user’s message. It includes the system prompt, any tool definitions, the conversation so far and any retrieved documents - often ten times the size of the question itself. The best source is the usage figures your provider returns with each response; average them over a day of real traffic.
Models that reason before answering may bill that hidden reasoning as output tokens. If yours does, include it in output tokens per request, or the estimate will be low.
Formula
Cost per request
(input tokens × input price + output tokens × output price) ÷ 1,000,000
Monthly cost
cost per request × requests per day × days per month
Cost per 1,000 requests
cost per request × 1,000
Blended cost per 1M tokens
cost per request ÷ (input + output tokens) × 1,000,000
Worked examples
A customer chat feature
10,000 requests a day for 30 days, 1,200 input and 350 output tokens at $3 and $15 per 1M: $0.0036 of input plus $0.00525 of output is $0.00885 per request - $88.50 a day, $2,655 a month and $31,860 a year. That is $8.85 per 1,000 requests and a blended $5.71 per 1M tokens, with output 59% of the cost.
An internal tool used on working days
500 requests a day for 22 days, 4,000 input and 1,000 output tokens at $1 and $5 per 1M: $0.009 per request, $4.50 a day, $99 a month and $1,188 a year. The blended price is $1.80 per 1M tokens.
Frequently asked questions
Why is output more expensive than input?
Input tokens are processed in parallel, while output is generated one token at a time, which uses the hardware for longer per token. Providers price that difference in, commonly at several times the input rate.
How many tokens is a typical request?
It varies widely: a short chat turn with a small system prompt may be a few hundred tokens, while a retrieval request with several documents can be tens of thousands. Measure your own from API usage data rather than guessing.
Does this include prompt caching or batch discounts?
No - it uses list prices. Use the cache savings and batch cost calculators to see how much those would take off.
Why is the annual figure monthly × 12?
So it matches your monthly budget exactly. If your traffic varies by season, run the calculator for a busy and a quiet month instead.
Are the preset prices up to date?
They were checked against the provider’s pricing page on the date shown with the presets. Prices change, so confirm them before relying on the result.
