AI & Machine Learning
AI Context Window Calculator
Calculate how much context your prompt, conversation and documents need, and whether they fit in the model’s context window.
Free to useNo sign-up requiredNo watermarkRuns in your browser
Last reviewed: 30 September 2026 by Vishal Senthilkumar
A model’s context window is a single budget shared by everything in the request - the system prompt, the conversation so far, any documents you retrieve or paste in - and by the answer it writes back. Run past it and the request fails, or older messages get cut without warning.
This calculator adds those parts up, keeps a safety reserve, and tells you whether the request fits, how much input you could still add, and the largest output that would fit. You can enter sizes in tokens, or in words or characters when that is all you have.
How this tool works
Set the context window
Pick a common size or type your model’s exact limit from its documentation.
Enter what goes in
System prompt, the user’s message, conversation history and any documents - in tokens, words or characters.
Enter the expected output
How long you expect the answer to be, in tokens. It uses the same window.
Read the verdict
Fits, fits only by using the reserve, or does not fit - with the numbers behind it.
How it works
Total input is the sum of the system prompt, the user prompt, the conversation history and the documents. When you enter words or characters, they are converted to tokens with English averages - about 1.33 tokens per word and 4 characters per token.
The safety reserve is a percentage of the window set aside for everything that is hard to predict: tokenizer differences, tool definitions and results, message formatting overhead, and a longer answer than planned.
Required context is input plus expected output plus reserve. What is left over is the window minus that. Room for more input keeps your expected output and reserve; the largest possible output assumes you add nothing more.
Common use cases
- Checking whether a retrieval-augmented (RAG) prompt still fits after adding more retrieved chunks.
- Deciding when a chat application should summarise or drop older messages.
- Choosing between a model with a larger window and splitting a document into pieces.
- Setting a sensible maximum output length for an API request.
Why token counts are only estimates
Each model family uses its own tokenizer, so the same text can be a different number of tokens in different models. English prose averages about three quarters of a word per token, but code, numbers, URLs, tables and most non-English languages use noticeably more tokens for the same length of text.
For a precise answer, count tokens with your provider’s tokenizer or token-counting endpoint and enter the exact figures here. The reserve exists because even exact counts do not include every piece of formatting overhead a provider adds.
Formula
Total input
system + user + conversation + documents
Reserve
context window × reserve %
Required context
total input + expected output + reserve
Left over
context window − required context
Largest possible output
context window − reserve − total input
Also limited by the model’s own maximum output setting.
Worked examples
A support chatbot with retrieval, 128K window
A 1,500-token system prompt, 500-token question, 6,000 tokens of history and 20,000 tokens of retrieved articles make 28,000 tokens of input. With a 4,000-token answer and a 10% reserve (12,800 tokens), 44,800 tokens are needed - leaving 83,200 spare.
Summarising a long report, 32K window
A 25,000-word report is roughly 33,000 tokens on its own - already more than a 32,768-token window before any prompt or answer. Split it into sections, summarise each, then combine the summaries.
Frequently asked questions
Does the output count towards the context window?
For most models, yes: the window is shared by the input and the generated output. Many models also have a separate, smaller limit on output length, so the largest output shown here may be higher than the model will actually produce.
How many tokens is a word?
For typical English text, about 1.3 tokens per word, or about 4 characters per token. It varies by tokenizer and by content - code and non-English text usually need more tokens. Use exact token counts where you can.
What reserve should I keep?
Around 10% is a sensible default for chat and retrieval applications, with at least a thousand tokens on small windows. Keep more if you use tools or function calling, whose definitions and results also take up context.
What happens if a request does not fit?
Depending on the provider, the request is rejected with a context-length error, or older messages are truncated. Either way, trim or summarise the conversation, retrieve fewer or shorter documents, or use a model with a larger window.
Is anything I type sent anywhere?
No. You enter numbers, not your prompt, and the calculation runs in your browser.
