Inference Latency Calculator
Estimate model inference latency, throughput, and memory requirements across CPU, GPU, and edge hardware.
About this calculator
This calculator gives a rough order-of-magnitude estimate of how long a single inference pass takes and how much memory a model needs, based on Model Parameters (millions), the target Hardware Type, Batch Size, and whether the model uses INT8 Quantization. Model Memory scales directly with parameter count and byte width: roughly 4 bytes per parameter at full precision (FP32) versus about 1 byte per parameter with INT8 quantization, so quantizing a model shrinks its memory footprint by roughly 4x. Inference Latency is modeled as whichever is larger of two costs -- the time to stream the model's weights from memory (set by the hardware's memory bandwidth) or the time to execute its floating-point operations (set by the hardware's compute throughput) -- because a real accelerator is idle waiting on whichever of those two resources is the bottleneck for the given model and batch size.
A fixed 1 ms floor is then added on top of that bottleneck cost to represent minimum kernel-launch and dispatch overhead; for very small models this floor dominates the reported latency, so the number you see there is mostly fixed overhead rather than memory or compute time. Raising Model Parameters always raises both Inference Latency and Model Memory, since weight-loading time, compute time, and memory footprint all scale directly with parameter count; Hardware Type and INT8 Quantization then rescale that base cost up or down without changing the direction of the relationship. Treat this output as an order-of-magnitude planning figure rather than a cycle-accurate hardware simulation: it uses fixed representative throughput and bandwidth figures for three generic hardware classes rather than a specific GPU/CPU SKU, ignores operator-level overhead, kernel launch latency, KV-cache growth for autoregressive generation, multi-GPU communication cost, and framework/runtime overhead -- all of which shift real measured latency away from this estimate, sometimes substantially.
Inputs
Results
Inference Latency
1.4 ms
Model Memory
400 MB
≈ 4 apps
How to Use This Calculator
- Enter Model Parameters (millions) for the model you're estimating.
- Select Hardware Type (CPU, GPU, or Edge) for the deployment target.
- Set Batch Size for the number of inputs processed simultaneously.
- Choose whether INT8 Quantization is enabled.
- Review Inference Latency (ms) and Model Memory (MB).
- Use Throughput (QPS) to inform your decision.
How the result changes with Model Parameters (millions)
| Model Parameters (millions) | Inference Latency | Model Memory |
|---|---|---|
| 50 | 1.2 ms | 200 MB |
| 75 | 1.3 ms | 300 MB |
| 150 | 1.7 ms | 600 MB |
| 250 | 2.1 ms | 1,000 MB |
What each input means
- Model Parameters (millions)
- Model size in millions of parameters
- Hardware Type
- Target deployment hardware — GPU offers highest throughput, edge for low-power
- Batch Size
- Number of inputs processed simultaneously (1 for real-time serving)
- INT8 Quantization
- INT8 quantization reduces model size ~4x and improves latency with minor accuracy loss
How this is calculated
Worked example, using the default values
- Identify Input Parameters4 parametersModel Parameters (millions) = 100, Hardware Type = 2, Batch Size = 1, INT8 Quantization = 0 = 4 input(s) provided
- Calculate Inference LatencyInference Latency1.44 = 1.44
- Calculate ThroughputThroughput694.4 = 694.4
Engine last updated . Checked against 2 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.
Frequently Asked Questions
What happens to Inference Latency and Model Memory as the model gets bigger?
Both rise. Raising Model Parameters (millions) always raises Model Memory (it scales directly with parameter count and byte width) and always raises Inference Latency across this calculator's full 0.1 to 1,000,000 range, since both the weight-loading cost and the compute cost it is built from grow with parameter count. Hardware Type and INT8 Quantization change how steep that climb is, but not its direction.
Why doesn't increasing Batch Size always increase Inference Latency in this calculator?
Because latency is modeled as whichever is larger of a memory-bound cost (loading the model's weights, which doesn't depend on batch size) and a compute-bound cost (running the arithmetic, which scales with Batch Size). At small batch sizes the memory-bound cost dominates and stays flat as Batch Size rises; once the compute-bound cost grows large enough to take over, Inference Latency starts climbing with Batch Size. It never decreases as Batch Size increases across this calculator's full 1 to 512 range.
Does the Hardware Type selection change Model Memory?
No -- Model Memory depends only on Model Parameters and whether INT8 Quantization is enabled, since it is purely a function of how many parameters exist and how many bytes each one occupies in memory. Switching between CPU, GPU, and Edge hardware changes Inference Latency and Throughput (via each hardware class's assumed compute and memory-bandwidth figures) but leaves Model Memory unchanged.
How much does INT8 Quantization actually save?
Quantization roughly quarters Model Memory, since it stores each parameter in about 1 byte instead of 4, and it also reduces the modeled compute cost contributing to Inference Latency, reflecting the real-world speedup INT8 arithmetic typically offers over FP32 on hardware that supports it. The tradeoff, noted in the field's own help text, is a minor accuracy loss versus the full-precision model.
Is this calculator's latency estimate accurate for a specific real GPU or CPU?
No -- treat the number as a ballpark for capacity planning, not a stand-in for a real benchmark run. It uses one representative throughput and memory-bandwidth figure per hardware class (CPU, GPU, Edge) rather than modeling a specific chip, and it ignores factors like kernel launch overhead, KV-cache growth, multi-GPU communication, and framework overhead that shift real measured latency -- sometimes substantially -- away from this simplified estimate.
Related Calculators
The questions that sit next to this one — chosen by subject, including calculators filed under a different category.
GPU Training Cost Calculator
Estimate cloud GPU training time, cost, and CO2 emissions for machine learning model training.
MLOps & AI CostingModel Compression Calculator
Calculate compressed model size after pruning, quantization, and distillation with accuracy impact estimates.
MLOps & AI CostingMLOps Pipeline Cost Calculator
Estimate monthly MLOps costs including training, serving, storage, and monitoring infrastructure.
Machine Learning & AIBayesian Inference Calculator
Calculate posterior probability, Bayes factor, likelihood ratio, and evidence strength using Bayes' theorem.
More in Technology & Computing.