Skip to main content
Calcimator

Gradient Descent Calculator

Calculate gradient descent parameters, convergence rate, effective learning rate, and training time estimates.

About this calculator

This calculator walks through the mechanics of mini-batch gradient descent training rather than modeling any specific neural network. Steps per Epoch divides Dataset Size by Batch Size, and Total Steps multiplies that by Number of Iterations (used here as an epoch count) -- Training Time is estimated from Total Steps at a fixed 0.1 seconds per compute step, so a larger Dataset Size or a smaller Batch Size (more steps per epoch) both increase the time estimate, while Learning Rate and Momentum do not, since neither changes how many gradient-update steps a run performs. Convergence Rate applies a simplified exponential-decay model, loss ∝ (1 − Learning Rate)^iterations -- NOT a real loss curve for any specific model, just an illustrative shape showing that a higher Learning Rate and more Number of Iterations both push this idealized "loss" toward zero faster.

Effective Learning Rate is Learning Rate / (1 − Momentum), the standard textbook approximation for how much momentum amplifies the step size a plain SGD update would take, growing sharply as Momentum approaches 1 (values above roughly 0.99 are rarely used in practice precisely because the amplification becomes destabilizing). Final Learning Rate applies a fixed 0.95-per-iteration exponential decay schedule -- a common but arbitrary choice among many possible learning-rate schedules, not a property of Learning Rate or Number of Iterations you set directly. The Simulated Loss Decay chart plots this same idealized (1 − Learning Rate)^epoch curve at 8 checkpoints, meant to illustrate the SHAPE of exponential convergence, not predict a real training run's actual loss values.

Inputs

Results

Steps per Epoch

313

Total Steps

313,000

Effective Learning Rate

0.1

Convergence Rate0
Training Time521.67 minutes
Final Learning Rate0
How to Use This Calculator
  1. Enter the learning rate (alpha), the number of training iterations, and the momentum coefficient.
  2. Set the batch size and dataset size to determine steps per epoch and total steps.
  3. Review the steps per epoch and total steps calculated for your training run.
  4. Check the effective learning rate (adjusted for momentum), convergence rate, and final learning rate after decay.
  5. Use the training time estimate and simulated loss decay chart to gauge how long training will take and how quickly loss decreases.

How the result changes with Batch Size

Batch SizeSteps per EpochTotal StepsEffective Learning Rate
16625625,0000.1
24417417,0000.1
48209209,0000.1
80125125,0000.1

What each input means

Learning Rate
Learning rate (step size)
Number of Iterations
Total training iterations
Momentum
Momentum coefficient
Batch Size
Samples per batch
Dataset Size
Total training samples

How this is calculated

Formula

θ_new = θ_old - α × ∇J(θ)

Worked example, using the default values

  1. Identify Input Parameters
    5 parameters
    Learning Rate = 0.01, Number of Iterations = 1000, Momentum = 0.9, Batch Size = 32, Dataset Size = 10000 = 5 input(s) provided
  2. Calculate Steps per Epoch
    Steps per Epoch
    313 = 313
  3. Calculate Total Steps
    Total Steps
    313000 = 313000
  4. Calculate Effective Learning Rate
    Effective Learning Rate
    0.1 = 0.1
  5. Calculate Convergence Rate
    Convergence Rate
    0 = 0
  6. Calculate Training Time
    Training Time
    521.67 = 521.67

Engine last updated . Checked against 2 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Why does Training Time depend on Dataset Size and Batch Size but not Learning Rate or Momentum?

Training Time is derived from Total Steps -- the number of gradient-update steps a run performs, which is set by Dataset Size, Batch Size, and Number of Iterations alone. Learning Rate and Momentum change HOW MUCH each step moves the model's parameters, not how many steps a training run takes, so they have no effect on this calculator's time estimate even though they strongly affect Convergence Rate.

Is Convergence Rate a real loss curve for my model?

No -- it's an idealized (1 − Learning Rate)^iterations exponential-decay shape, useful for building intuition about how Learning Rate and Number of Iterations trade off, but it isn't derived from any specific model, dataset, or loss function. Real training loss curves are noisy, can plateau or spike, and depend on the model architecture and data in ways this simplified formula does not capture.

Why does Effective Learning Rate grow so quickly as Momentum approaches 1?

Effective Learning Rate is Learning Rate divided by (1 − Momentum), and that denominator shrinks toward zero as Momentum climbs toward 1 -- at Momentum = 0.9 the divisor is 0.1 (a 10x amplification), but at Momentum = 0.99 it's 0.01 (a 100x amplification). This is why momentum values above roughly 0.9-0.95 are used cautiously in practice: the effective step size can grow large enough to destabilize training.

What does the fixed 0.95 decay rate in Final Learning Rate represent?

It's one specific, commonly-used learning-rate decay schedule -- multiplying the rate by 0.95 at every iteration -- not a value derived from Learning Rate or Number of Iterations you enter. Real training pipelines use many different decay schedules (step decay, cosine annealing, warmup-then-decay); this calculator fixes one illustrative schedule so Final Learning Rate has a concrete number to report.

Why do Total Steps and Steps per Epoch matter if I already set Number of Iterations?

Number of Iterations sets how many passes (epochs) through the dataset a run performs, but each epoch itself takes multiple gradient-update steps -- one per mini-batch -- determined by Dataset Size divided by Batch Size. Total Steps (Steps per Epoch × Number of Iterations) is the figure that actually drives Training Time, since real wall-clock cost tracks individual gradient updates, not epochs.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.