VSThiran

AI & Machine Learning

AI Prompt Cache Savings Calculator

Work out how much prompt caching cuts your AI input bill, including cache-write costs and cache misses.

Free to useNo sign-up requiredNo watermarkRuns in your browser

Last reviewed: 30 September 2026 by Vishal Senthilkumar

Most production prompts repeat themselves. The same system prompt, the same tool definitions, the same policy document or codebase context is sent again on every request, and without caching you pay full input price for all of it every time.

Prompt caching lets the provider reuse that repeated prefix. This calculator compares your input bill with and without caching, counts the extra you pay when the cache has to be written, and shows the hit rate below which caching stops paying for itself.

How this tool works

  1. Enter the prompt sizes

    Total input tokens per request, and how many of them are an identical prefix that repeats on every call.

  2. Set a realistic hit rate

    Steady traffic keeps the cache warm; bursty or low traffic means more misses and more cache writes.

  3. Enter your prices

    Pick a preset or type your provider’s input, cache-read and cache-write prices per million tokens.

  4. Compare the totals

    Daily, monthly and annual cost with and without caching, the saving, and the break-even hit rate.

How it works

A cache stores the model’s processed state for the start of a prompt. When a later request begins with exactly the same tokens, that part is read from the cache instead of being processed again, and is billed at the much lower cache-read price. Anything after the first changed token is billed normally.

The first request, and any request that arrives after the cache has expired, has to write the prefix to the cache. Some providers charge a premium for this (for example 1.25× the input price for a short-lived cache and more for a longer one); others charge only the normal input price. The hit rate decides how often you pay the read price and how often the write price.

For each request, the uncached part of the prompt is priced at the normal input rate and the cached prefix at a blended rate: hit rate × read price + miss rate × write price. Multiplying by requests per day and days per month gives the monthly totals. Output tokens cost the same either way, so they are left out.

The break-even hit rate is where the blended rate equals the normal input price. If your provider charges no write premium, caching never costs more and the break-even is zero.

Common use cases

  • Chatbots and assistants with a long system prompt and tool definitions sent on every turn.
  • Question answering over the same handbook, contract or codebase for many users.
  • Multi-turn conversations, where each turn re-sends the whole history as a growing prefix.
  • Agent loops that call the model many times with the same instructions and tool list.
  • Deciding whether a longer-lived (more expensive to write) cache is worth it for bursty traffic.

How to get a high cache hit rate

Caches match prefixes, so order matters: put everything static first (system prompt, tool definitions, reference documents) and everything that changes last (the user’s message, timestamps, retrieved snippets). A single changed token near the start - a date in the system prompt, a reordered tool list - invalidates everything after it.

Caches also expire. Lifetimes are typically a few minutes, refreshed each time the prefix is read, with longer options on some providers at a higher write price. Steady traffic keeps a cache warm; a request every half hour may miss every time. Most providers also have a minimum cacheable length, commonly around a thousand tokens.

Formula

Blended cached rate

hit rate × cache-read price + (1 − hit rate) × cache-write price

Cost per request with caching

(input − cached prefix) × input price + cached prefix × blended rate

Prices per 1M tokens, so divide by 1,000,000.

Monthly cost

cost per request × requests per day × days per month

Break-even hit rate

(write price − input price) ÷ (write price − read price)

Worked examples

A support assistant with a large shared prefix

20,000 input tokens per request, 15,000 of them a fixed prefix, 5,000 requests a day for 30 days, at $3 input, $0.30 cache read and $3.75 cache write per 1M. At a 95% hit rate the blended cached rate is $0.4725 per 1M, so input drops from $9,000 to $3,313.13 a month - a saving of $5,686.88 (63.2%), or $68,242.50 a year. Caching breaks even at about a 21.7% hit rate.

When caching costs more

The same prices with a 10,000-token prompt cached in full but only a 10% hit rate: the blended rate is $3.405 per 1M, above the $3 input price. 150,000 requests a month cost $5,107.50 instead of $4,500 - caching loses $607.50 because nearly every request pays the write premium.

Frequently asked questions

Does prompt caching reduce output costs?

No. Caching only changes what you pay for input tokens. Output is generated fresh every time and billed at the normal output price, which is why this calculator leaves it out.

Why does a cache write cost more than a normal input token?

Storing the processed prefix uses memory on the provider’s side, and some providers pass that on as a write premium. If yours charges no premium, enter the normal input price as the write price - caching then never costs more than not caching.

What hit rate should I assume?

Measure it if you can: most APIs report cached and uncached input tokens in each response’s usage data. As a starting point, busy endpoints with a stable prefix often see very high hit rates, while low-traffic or bursty workloads see many misses as the cache expires between requests.

Can I combine caching with batch discounts?

With several providers, yes - the discounts stack. Batch jobs can still miss the cache, though, because requests in a batch are not guaranteed to run close together. Check your provider’s documentation.

Are the preset prices current?

They were copied from the provider’s published pricing page on the date shown under the presets. Prices change, so check the page before relying on them. You can overwrite every field.