GPU Training Cost Calculator
Estimate cloud GPU training time, cost, and CO2 emissions for machine learning model training.
About this calculator
This calculator estimates cloud GPU training cost using a simplified compute-scaling model rather than a live cloud-pricing feed. Training compute follows the widely-cited approximation FLOPs per step ≈ 6 × parameters × batch size (from Kaplan et al.'s neural scaling-law work), and the calculator assumes a dataset of roughly 1 token per model parameter -- a deliberate simplification, not the ~20-tokens-per-parameter "Chinchilla-optimal" ratio many production runs target -- to convert Model Parameters (millions) and Training Epochs into total FLOPs. Estimated Training Time (hrs) divides that FLOP count by GPU Type's rated TFLOPS at an assumed 35% model FLOPs utilization (MFU), the middle of the 30-40% range commonly achieved in real large-model training runs and well below each GPU's theoretical peak throughput. Total Training Cost multiplies Estimated Training Time by Cost per Hour, and the two move in OPPOSITE directions as GPU Type changes: a faster, pricier GPU (H100) finishes sooner, so Total Training Cost does not simply track the hourly price.
Doubling Model Parameters (millions), by contrast, roughly QUADRUPLES total FLOPs -- under this calculator's own dataset-size assumption, a larger model also implies a proportionally larger dataset -- so Model Parameters (millions) moves Total Training Cost steeply, on top of whatever GPU Type contributes through its own price and speed. Batch Size is deliberately NOT one of Total Training Cost's inputs: this is the real Kaplan/Chinchilla result that total training compute (FLOPs per step times steps per epoch) is batch-size invariant, so changing Batch Size alone leaves Estimated Training Time, Total Training Cost, and CO₂ Emissions unchanged -- it only changes how that same total compute is split across individual training steps, which this simplified model doesn't estimate separately. CO₂ Emissions applies each GPU's rated TDP (thermal design power) times a 1.3× power-usage-effectiveness (PUE) overhead for datacenter cooling and infrastructure, then a rough 0.5 kg CO₂-per-kWh grid-average factor -- both figures are commonly-cited round numbers, not a location- or provider-specific carbon accounting.
Inputs
Results
Estimated Training Time
0.5 hrs
Total Training Cost
$1.69
How to Use This Calculator
- Enter Model Parameters (millions) and Training Epochs — these drive total training compute.
- Set Batch Size (affects memory and gradient noise, but not this calculator's time/cost/CO2 estimates — see the FAQ) and select a GPU Type: NVIDIA T4 ($0.76/hr), NVIDIA A100 ($3.67/hr), or NVIDIA H100 ($8.50/hr).
- Review Estimated Training Time (hrs) and Total Training Cost ($).
- Use Cost per Hour ($) and CO₂ Emissions (kg) to inform your decision.
How the result changes with Model Parameters (millions)
| Model Parameters (millions) | Estimated Training Time | Total Training Cost |
|---|---|---|
| 50 | 0.1 hrs | $0.40 |
| 75 | 0.3 hrs | $0.95 |
| 150 | 1 hrs | $3.78 |
| 250 | 2.9 hrs | $10.50 |
What each input means
- Model Parameters (millions)
- Number of model parameters in millions (e.g., 125M for GPT-2 small)
- Training Epochs
- Number of complete passes through the training dataset
- Batch Size
- Number of samples processed per training step
- GPU Type
- Cloud GPU instance type — prices based on typical on-demand rates
How this is calculated
Worked example, using the default values
- Identify Input Parameters4 parametersModel Parameters (millions) = 100, Training Epochs = 3, Batch Size = 32, GPU Type = 2 = 4 input(s) provided
- Calculate Estimated Training TimeEstimated Training Time0.46 = 0.46
- Calculate Total Training CostTotal Training Cost1.69 = $1.69
- Calculate Cost per Hour3.67 = $3.67
Engine last updated . Checked against 2 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.
Frequently Asked Questions
Why doesn't switching to a faster, more expensive GPU always raise Total Training Cost?
Estimated Training Time and Cost per Hour move in opposite directions as GPU Type changes. An NVIDIA H100 costs more per hour than a T4, but its far higher TFLOPS rating finishes the same total FLOPs in a fraction of the time, so at this calculator's default settings an H100 run's Total Training Cost comes out LOWER than the same job on a T4 or A100 -- raw hourly price alone doesn't determine which GPU is cheapest for a given job.
Why does doubling Model Parameters (millions) roughly quadruple Total Training Cost instead of doubling it?
This calculator assumes dataset size scales with model size, roughly 1M tokens per 1M parameters, rather than holding dataset size fixed. Doubling Model Parameters (millions) therefore doubles BOTH the FLOPs needed per training step AND the number of steps per epoch, so total compute -- and Total Training Cost -- scales closer to the square of Model Parameters (millions), not linearly with it.
Where do the T4/A100/H100 per-hour prices come from, and how current are they?
They represent typical on-demand cloud GPU rental rates at the time this calculator was built, meant to be broadly representative rather than a live feed from any specific vendor. Actual prices vary by provider, region, spot vs. on-demand terms, and change over time, sometimes substantially -- treat Total Training Cost as an order-of-magnitude planning estimate, not a quote.
What does the 35% GPU utilization assumption mean, and why not 100%?
Real large-model training essentially never reaches a GPU's advertised peak TFLOPS -- data loading, communication between GPUs, memory bandwidth limits, and kernel inefficiencies mean most published large-scale training runs report a Model FLOPs Utilization (MFU) in roughly the 30-40% range. This calculator fixes that at 35%, the middle of that commonly-reported range, so results represent a realistic training run rather than an idealized best case.
Why is CO₂ Emissions only a rough estimate rather than a precise carbon-footprint figure?
It combines each GPU's rated thermal design power (TDP) with a flat 1.3× power-usage-effectiveness overhead for datacenter cooling and infrastructure, and a single ~0.5 kg CO₂-per-kWh grid-average factor -- both commonly-cited round numbers, not measurements tied to a specific datacenter, region, or electricity mix. A real facility's PUE and grid carbon intensity can each vary by a factor of two or more depending on location and renewable energy mix.
Related Calculators
The questions that sit next to this one — chosen by subject, including calculators filed under a different category.
Inference Latency Calculator
Estimate model inference latency, throughput, and memory requirements across CPU, GPU, and edge hardware.
MLOps & AI CostingMLOps Pipeline Cost Calculator
Estimate monthly MLOps costs including training, serving, storage, and monitoring infrastructure.
MLOps & AI CostingTraining Data Size Calculator
Estimate minimum dataset size needed for machine learning models based on features, complexity, and accuracy targets.
More in Technology & Computing.