VSThiran

AI & Machine Learning

AI Inference Cost Calculator

Turn token counts and per-million prices into cost per request, per month and per year for an LLM API.

Free to useNo sign-up requiredNo watermarkRuns in your browser

Last reviewed: 30 September 2026 by Vishal Senthilkumar

LLM APIs charge per token, with separate prices for the tokens you send and the tokens the model writes - and output usually costs several times more than input. That makes it hard to see, from a price table alone, what a feature will cost once real users are on it.

Enter your traffic and the typical size of a request and response, and this calculator gives the cost per request, per day, per month and per year, the cost per 1,000 requests, and a single blended price per million tokens for your particular input-output mix.

How this tool works

  1. Enter your traffic

    Requests per day and how many days a month the feature runs.

  2. Enter typical token counts

    Input tokens per request (including the system prompt and context) and output tokens per response.

  3. Enter prices

    Input and output prices per 1M tokens, or pick a preset.

  4. Read the costs

    Per request, per day, per month, per year, per 1,000 requests and blended per 1M tokens.

How it works

The cost of one request is input tokens × input price ÷ 1,000,000 plus output tokens × output price ÷ 1,000,000. Daily cost multiplies that by requests per day, monthly by days per month, and annual cost is twelve months.

The blended cost per 1M tokens divides the cost of a request by all the tokens in it. Because output is priced higher, a workload with long answers has a blended price close to the output price; a workload with long prompts and short answers sits close to the input price.

The share of cost from output tells you which lever matters more. If output dominates, shorter answers and a sensible maximum output length help most; if input dominates, trimming context and prompt caching do.

Common use cases

  • Budgeting a new chatbot, assistant or AI feature before launch.
  • Comparing two models by entering each one’s prices against the same traffic.
  • Setting a price or usage limit for an AI feature in your own product.
  • Checking whether a cheaper model or shorter prompts are worth the effort.

Getting realistic token counts

Input is more than the user’s message. It includes the system prompt, any tool definitions, the conversation so far and any retrieved documents - often ten times the size of the question itself. The best source is the usage figures your provider returns with each response; average them over a day of real traffic.

Models that reason before answering may bill that hidden reasoning as output tokens. If yours does, include it in output tokens per request, or the estimate will be low.

Formula

Cost per request

(input tokens × input price + output tokens × output price) ÷ 1,000,000

Monthly cost

cost per request × requests per day × days per month

Cost per 1,000 requests

cost per request × 1,000

Blended cost per 1M tokens

cost per request ÷ (input + output tokens) × 1,000,000

Worked examples

A customer chat feature

10,000 requests a day for 30 days, 1,200 input and 350 output tokens at $3 and $15 per 1M: $0.0036 of input plus $0.00525 of output is $0.00885 per request - $88.50 a day, $2,655 a month and $31,860 a year. That is $8.85 per 1,000 requests and a blended $5.71 per 1M tokens, with output 59% of the cost.

An internal tool used on working days

500 requests a day for 22 days, 4,000 input and 1,000 output tokens at $1 and $5 per 1M: $0.009 per request, $4.50 a day, $99 a month and $1,188 a year. The blended price is $1.80 per 1M tokens.

Frequently asked questions

Why is output more expensive than input?

Input tokens are processed in parallel, while output is generated one token at a time, which uses the hardware for longer per token. Providers price that difference in, commonly at several times the input rate.

How many tokens is a typical request?

It varies widely: a short chat turn with a small system prompt may be a few hundred tokens, while a retrieval request with several documents can be tens of thousands. Measure your own from API usage data rather than guessing.

Does this include prompt caching or batch discounts?

No - it uses list prices. Use the cache savings and batch cost calculators to see how much those would take off.

Why is the annual figure monthly × 12?

So it matches your monthly budget exactly. If your traffic varies by season, run the calculator for a busy and a quiet month instead.

Are the preset prices up to date?

They were checked against the provider’s pricing page on the date shown with the presets. Prices change, so confirm them before relying on the result.