Skip to main content
Calcimator

Inference Latency Calculator

Estimate model inference latency, throughput, and memory requirements across CPU, GPU, and edge hardware.

Inputs

Results

Inference Latency

1.4 ms

Model Memory

400 MB

≈ 4 apps

Throughput694.4 QPS
How to Use This Calculator
  1. Enter Model Parameters (millions), Hardware Type, and CPU (x86).
  2. Set GPU (A100), Edge (ARM), and Batch Size.
  3. Adjust INT8 Quantization, No (FP32) as needed.
  4. Review Inference Latency (ms) and Model Memory (MB).
  5. Use Throughput (QPS) to inform your decision.

How the result changes with Model Parameters (millions)

Model Parameters (millions)Inference LatencyModel Memory
100,000445.4 ms400,000 MB
350,0001,556.6 ms1,400,000 MB
650,0002,889.9 ms2,600,000 MB
900,0004,001 ms3,600,000 MB

What each input means

Model Parameters (millions)
Model size in millions of parameters
Hardware Type
Target deployment hardware — GPU offers highest throughput, edge for low-power
Batch Size
Number of inputs processed simultaneously (1 for real-time serving)
INT8 Quantization
INT8 quantization reduces model size ~4x and improves latency with minor accuracy loss

How this is calculated

Worked example, using the default values

  1. Identify Input Parameters
    4 parameters
    Model Parameters (millions) = 100, Hardware Type = 2, Batch Size = 1, INT8 Quantization = 0 = 4 input(s) provided
  2. Calculate Inference Latency
    Inference Latency
    1.44 = 1.44
  3. Calculate Throughput
    Throughput
    694.4 = 694.4

Engine last updated . Checked against 2 independently-derived tests how we verify calculators.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.