Skip to main content
Calcimator

LLM Token Calculator

Estimate token count and API cost from text length across different tokenizers (GPT-4, Claude, Llama).

About this calculator

Different LLM providers split text into tokens differently, and that difference matters for both cost and context-window planning. This calculator estimates token count with a simple characters-per-token ratio tuned per tokenizer family — about 4.0 for GPT-4's cl100k_base, 3.8 for GPT-3.5's older p50k_base, 3.5 for Claude, and 3.2 for Llama's SentencePiece — dividing your character count by that ratio and rounding up. It's a real approximation, not an exact tokenizer run: actual token counts depend on the specific words, punctuation, and whitespace in your text, so treat the result as a solid estimate for budgeting rather than an exact count. Output tokens are derived from input tokens using an output-to-input ratio you set, since output length varies hugely by task — summarization typically runs far shorter than the input (around 0.3x), while open-ended generation can run 2x or more.

Cost is computed by pricing input tokens at your entered rate and pricing output tokens at three times that rate, reflecting the fact that most providers charge noticeably more for generated tokens than for tokens they merely read. The total is then multiplied across your full request count for a batch cost estimate. The context-usage output shows what percentage of a 128K-token context window a single request would consume — useful for catching prompts that are approaching a hard context limit before they fail in production. The most common mistake is entering price-per-token instead of price-per-million-tokens, which will produce a wildly wrong cost by six orders of magnitude.

Inputs

Results

Input tokens per request

1,250

Total API cost ($)

$0.02

Output tokens per request1,250
Total tokens per request2,500
Total tokens (all requests)2,500
Input cost ($)$0.00
Output cost ($)$0.01
128K context usage (%)1.95%
Tokenizer NameGPT-4 (cl100k_base)
How to Use This Calculator
  1. Enter the Text Length in characters — paste your typical prompt and count characters, or use a rough estimate.
  2. Select the Tokenizer matching your model (GPT-4 uses ~4 chars/token, Claude ~3.5 chars/token).
  3. Set the Input Price per 1M Tokens from the provider's pricing page and the Output-to-Input Ratio (e.g., 0.3 for summarization, 2 for generation).
  4. Enter the Number of Requests to estimate total batch cost.
  5. Review Input Tokens, Output Tokens, Total Cost, and Context Usage % to optimize your prompt length and budget.

How the result changes with Text length (characters)

Text length (characters)Input tokens per requestTotal API cost ($)
2,500625$0.01
3,750938$0.01
7,5001,875$0.02
12,5003,125$0.04

What each input means

Text length (characters)
Number of characters in your input text. Average English word is ~5 characters.
Tokenizer
Select the tokenizer
Input price per 1M tokens ($)
Cost per 1 million input tokens. GPT-4o: $2.50, Claude Sonnet: $3, GPT-4 Turbo: $10.
Output-to-input ratio
Expected output tokens as a multiple of input tokens. Summarization ~0.3, Q&A ~1, generation ~2+.
Number of requests
Total API requests to estimate cost for.

What each result means

Input tokens per request
Estimated input token count per request.
Output tokens per request
Estimated output token count based on the output ratio.
Total tokens per request
Input + output tokens consumed per request.
Total tokens (all requests)
Grand total token consumption across all requests.
Input cost ($)
Cost for input tokens across all requests.
Output cost ($)
Cost for output tokens (typically 3x input price).
Total API cost ($)
Combined input + output token cost.
128K context usage (%)
Percentage of a 128K context window used per request.

How this is calculated

Worked example, using the default values

  1. Identify Input Parameters
    4 parameters
    Text length (characters) = 5000, Tokenizer = 0, Input price per 1M tokens ($) = 3, Output-to-input ratio = 1 = 5 input(s) provided
  2. Calculate Input tokens per request
    1250 = 1250
  3. Calculate Total API cost
    Total API cost = inputCost + outputCost
    0.015 = $0.015
  4. Calculate Output tokens per request
    Output tokens per request = ceil(inputTokens * outputRatio)
    1250 = 1250
  5. Calculate Total tokens per request
    Total tokens per request = inputTokens + outputTokens
    2500 = 2500

Engine last updated . Checked against 3 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Why does the tokenizer choice change the token count for the same text?

Each tokenizer family builds its vocabulary differently, so the same sentence splits into a different number of chunks depending on which one processes it. This calculator approximates that with a fixed characters-per-token ratio per tokenizer — 4.0 for GPT-4's cl100k_base down to 3.2 for Llama's SentencePiece — so switching the tokenizer at a fixed character count directly rescales the estimated input tokens; a smaller ratio means more tokens for the same text.

Why is output priced at three times the input rate instead of using a separate output price?

The calculator only takes one price input (per 1M input tokens) and multiplies it by 3 to approximate the output rate, reflecting that most commercial LLM APIs do charge noticeably more for generated tokens than for tokens they merely read. It's a simplification, not a lookup of your specific provider's real output price — if your provider's output-to-input price ratio differs from 3x, the outputCost figure here will be off by that same factor.

What does the 128K context usage percentage actually measure?

It divides your estimated total tokens per request (input plus output) by a fixed 128,000-token context window and expresses that as a percentage — it does not check your actual model's context limit, which may be smaller (like 8K or 32K) or larger. Use it as a rough proxy for how close a single request is to hitting a typical modern context ceiling, and adjust manually if your model's real window differs from 128K.

How should I set the output-to-input ratio for my task?

The ratio scales output tokens directly off input tokens, so it should reflect how much your task typically expands or shrinks text: around 0.3 for summarization (output much shorter than input), close to 1 for question-answering, and 2 or higher for open-ended generation where the model writes substantially more than it reads. Getting this wrong skews both outputCost and totalTokensPerRequest since output tokens are derived entirely from this multiplier rather than measured independently.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.