Fine-Tuning Cost Calculator
Estimate LLM fine-tuning costs from dataset size, epochs, and model choice for OpenAI API and self-hosted options.
About this calculator
Fine-tuning cost splits into two entirely different pricing models depending on where the training runs, and this calculator switches between them based on which model you pick. For the API-based options — GPT-4o mini, GPT-4o, and GPT-3.5 Turbo — the tool separates your dataset into a training portion and a validation portion by the split percentage, but both portions still get priced at the same per-million-token training rate and summed into one bill: the split changes how the total is reported (a "validation examples" count shows up as its own line), but it does not change the total number of tokens billed, since every example — training or validation — passes through training at least once per epoch. In other words, raising the validation split reshuffles which examples are labeled "training" versus "validation" without lowering API cost; GPT-4o's $25-per-million rate versus GPT-4o mini's $3 is what actually swings the total. The self-hosted options (Llama 3, Mistral) work differently: GPU-hours are estimated only from the training portion — about one A100 GPU-hour per 10,000 training examples per epoch for Llama 3, slightly less for Mistral — at an assumed $1.50/hour spot price, so a larger validation split does reduce estimated self-hosted compute cost, unlike the API path.
On top of raw training cost, the calculator adds a dataset preparation estimate — roughly 2 minutes of labeling time per example at $50/hour — which for smaller datasets often exceeds the training cost itself. A fixed 40% prompt-token savings figure is also shown, reflecting the common (but not guaranteed) benefit that a fine-tuned model needs a shorter prompt at inference. Confirm current rates against your provider's live pricing page before committing budget.
Inputs
Results
Total estimated cost ($)
$1,671.17
≈ 13 pairs of sneakers
How to Use This Calculator
- Enter the number of Training Examples (prompt-completion pairs) in your dataset — typically 500–5,000 for most fine-tuning tasks.
- Set average Tokens per Example (prompt + completion combined) and the number of Training Epochs (2–4 is typical).
- Select the Model — API-based fine-tuning (GPT-4o mini, GPT-3.5) or self-hosted (Llama, Mistral).
- Set the Validation Split percentage to hold out data for evaluation.
- Review Total Estimated Cost, Training Tokens, Estimated Training Time, and Dataset Prep Hours to plan your fine-tuning project.
How the result changes with Training examples
| Training examples | Total estimated cost ($) |
|---|---|
| 500 | $835.58 |
| 750 | $1,253.38 |
| 1,500 | $2,506.75 |
| 2,500 | $4,177.92 |
What each input means
- Training examples
- Number of training examples (prompt-completion pairs) in your dataset.
- Avg tokens per example
- Average total tokens (prompt + completion) per training example.
- Training epochs
- Number of passes through the training data. OpenAI recommends 2-4 for most tasks.
- Model
- Select the model
- Validation split (%)
- Percentage of examples reserved for validation. Typical: 10-20%.
What each result means
- Total estimated cost ($)
- Training cost + dataset preparation cost.
- Training/compute cost ($)
- API fine-tuning fee or self-hosted GPU compute cost.
- Dataset prep cost ($)
- Estimated data labeling cost at $50/hr (~2 min per example).
- Training tokens (millions)
- Total tokens processed during training across all epochs.
- Est. training time (hours)
- Estimated wall-clock training time.
- Dataset prep hours
- Estimated hours for manual data preparation and labeling.
- Validation examples
- Number of examples held out for validation.
- Prompt token savings (%)
- Estimated inference prompt savings vs. few-shot prompting (~40%).
How this is calculated
Worked example, using the default values
- Identify Input Parameters4 parametersTraining examples = 1000, Avg tokens per example = 500, Training epochs = 3, Model = 0 = 5 input(s) provided
- Calculate Total estimated costTotal estimated cost = trainingCost + datasetPrepCost1671.17 = $1,671.17
- Calculate Training/compute costTraining/compute cost4.5 = $4.5
- Calculate Dataset prep costDataset prep cost = datasetPrepHours * 501666.67 = $1,666.67
Engine last updated . Checked against 3 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.
Frequently Asked Questions
Does raising the Validation Split percentage lower my fine-tuning cost?
It depends entirely on which model you pick. For the three API-based models, no — every example, training or validation, still passes through training at least once per epoch, so the split only relabels which examples are reported as 'validation' without reducing total billed tokens. For the two self-hosted models, yes — GPU-hours are estimated only from the training portion, so a larger validation split does shrink estimated compute cost.
Why is Dataset Prep Cost sometimes bigger than Training/Compute Cost?
Dataset prep is estimated at a flat 2 minutes of labeling time per example at $50/hour, which scales linearly with your example count regardless of model. Training cost for cheap models like GPT-4o mini ($3 per million tokens) can be very small for modest datasets, so for smaller or lower-token datasets the fixed labeling estimate frequently outweighs the actual training bill — it's worth checking both lines separately rather than assuming training dominates.
How does switching from GPT-4o mini to GPT-4o change the results?
It changes the per-million-token training price used for both the training and validation portions from $3 to $25 — over 8x higher — which directly multiplies trainingCost and totalCost by that same factor for the same token volume. It does not change totalTrainingTokensM, datasetPrepCost, or promptSavingsPct, since none of those depend on which API model you selected.
What does the 40% prompt token savings figure represent, and is it guaranteed?
It's a fixed assumption baked into the calculator, not something computed from your inputs — it reflects the commonly observed pattern that a fine-tuned model needs a shorter prompt at inference time than a few-shot-prompted base model, since the fine-tuning has already taught it the desired behavior. It stays at 40% regardless of your dataset size, model choice, or epochs, so treat it as a general planning heuristic rather than a measured outcome for your specific task.
Related Calculators
The questions that sit next to this one — chosen by subject, including calculators filed under a different category.
LLM Token Calculator
Estimate token count and API cost from text length across different tokenizers (GPT-4, Claude, Llama).
Ai ToolsAI Model Cost Comparison Calculator
Compare per-token and per-request costs across AI providers including GPT-4o, Claude, Gemini, and self-hosted Llama.
Ai ToolsEmbedding Cost Calculator
Calculate embedding API costs and storage requirements across OpenAI, Cohere, and self-hosted models.
More in Technology & Computing.