AI & Machine Learning
AI Batch API Cost Calculator
Price a large AI job through a batch API and see what it saves over sending the same requests in real time.
Free to useNo sign-up requiredNo watermarkRuns in your browser
Last reviewed: 30 September 2026 by Vishal Senthilkumar
Plenty of AI work does not need an answer in seconds: classifying a backlog of tickets, summarising an archive, tagging a product catalogue, running an evaluation set. Batch APIs take that kind of job as a file of requests, process it within a set window, and charge less for it - often half the normal price.
Enter the size of the job and your prices to see the real-time cost, the batch cost, what you save, the cost per request and how many batch files the job splits into.
How this tool works
Describe the job
How many requests, and the average input and output tokens per request.
Enter prices
Standard input and output prices per 1M tokens, or pick a preset.
Set the discount and batch size
Your provider’s batch discount, and how many requests go in each batch file.
Read the comparison
Batch cost, real-time cost, savings, cost per request and number of batches.
How it works
The standard cost of one request is its input tokens times the input price plus its output tokens times the output price, with prices per million tokens. Multiplying by the number of requests gives what the job would cost sent one request at a time.
The batch cost applies your discount to that total. Where the discount covers both input and output - as it does on several major APIs - the saving is simply the discount percentage of the whole bill.
Providers cap how many requests (and how many megabytes) one batch file may hold, so a large job is split into several batches. The number of batches is the total requests divided by requests per batch, rounded up.
Common use cases
- Classifying, tagging or extracting fields from a large backlog of documents or tickets.
- Generating product descriptions, translations or summaries for a whole catalogue.
- Running evaluation sets and regression tests against a new prompt or model.
- Creating embeddings or synthetic training data overnight.
- Nightly or weekly reports where nobody is waiting on each individual answer.
When batching is the wrong choice
Batches are asynchronous. Results arrive when the provider finishes the job - often well within the window, but with no promise of seconds - so anything a user is waiting on, and any workflow where the next step depends on the previous answer, should stay real-time.
Batches also make failures slower to notice. Send a small test batch first to catch a malformed request or a prompt that produces the wrong format before you pay for a hundred thousand of them.
Formula
Standard cost per request
(input tokens × input price + output tokens × output price) ÷ 1,000,000
Batch cost
standard cost × (1 − batch discount)
Savings
standard cost − batch cost
Number of batches
⌈requests ÷ requests per batch⌉
Worked examples
Classifying 100,000 support tickets
1,500 input and 400 output tokens each at $3 and $15 per 1M cost $0.0105 per request, or $1,050 in real time. With a 50% batch discount the job costs $525 - $0.00525 per request - and at 10,000 requests per file it runs as 10 batches.
A job that does not divide evenly
120,001 requests of 1,000 input and 1,000 output tokens at $1 and $2 per 1M cost $360.00 in real time. A 30% discount brings it to $252.00, saving $108.00. At 50,000 requests per batch that is 3 batches - the last one holds a single request.
Frequently asked questions
How much cheaper is a batch API?
Several major providers offer 50% off both input and output tokens for batch jobs, but it varies by provider and model and can change. Enter the discount from your provider’s pricing page.
How long do batch jobs take?
Providers set a completion window - commonly up to 24 hours - and many batches finish much sooner. Plan as if you will get results at the end of the window.
How many requests can one batch hold?
It depends on the provider: there is usually a limit on requests per batch and on the file size. For example, OpenAI documents 50,000 requests and 200 MB per batch file. Set requests per batch to your provider’s limit or lower.
Can I use prompt caching in a batch?
Often, yes, and the discounts can stack. Because requests in a batch may not run close together, cache hits are less predictable than with steady real-time traffic.
Does the calculator include failed requests?
No. It prices every request once. If some fail and need resubmitting, add them to the request count.
