Skip to main content
Calcimator

Neural Network Trainer Calculator

Complete neural network training analysis. Architecture design, learning rates, batch sizes, regularization, and activation functions.

About this calculator

This calculator has five independent analysis modes, selected via the Analysis Type dropdown: Network Architecture (the default), Learning Rate Schedule, Batch Size & Training, Regularization, and Activation Functions -- each mode shows only its own relevant inputs and computes an entirely separate set of outputs. In the default Architecture mode, Total Parameters sums the weights and biases of every layer: each layer's parameter count is (inputs to that layer) x (outputs from that layer) + (outputs from that layer, for the bias terms). Which input matters most depends on the shape of the network, not on a fixed ranking. Hidden Units per Layer enters quadratically in every hidden-to-hidden matrix, but at the default 784-input architecture the FIRST layer alone holds about three-quarters of the parameters, so Input Size is the dominant term there; widen the input to 3,072 and it holds 92 percent. Hidden Units takes over only once the network is deep or narrow-input enough for the hidden-to-hidden matrices to outweigh the input projection. Hidden Layers similarly increases Total Parameters, since each additional layer adds its own full hidden-to-hidden weight matrix.

Batch Size has no effect on Total Parameters at all -- it only affects Total Memory, via the activation memory needed during training, not the model's fixed parameter count. Total Memory counts four fp32 buffers per parameter -- the weights, their gradients, and Adam's two moment estimates -- plus one activation value per neuron per sample in the batch. Plain SGD without momentum needs only two of those four, so it uses roughly half as much; mixed-precision training changes the arithmetic again. Complexity Rating classifies the resulting Total Parameters into Simple, Medium, Large, or "Very Large (LLM scale)" bands. The other four modes are independent tools sharing this same calculator: switching Analysis Type entirely changes which inputs are shown and which formulas run. The Regularization mode's strength score, the batch-size generalization-risk bands and the relative activation compute costs are rules of thumb chosen for this calculator, not measured or published figures -- use them to compare two configurations against each other, not as absolute numbers.

Progress0%

Step 1 of 2

How to Use This Calculator
  1. Pick an Analysis Type first — the five modes are independent tools, and each one shows only its own inputs and results.
  2. Network Architecture: enter Input Size, Hidden Layers, Hidden Units per Layer, Output Size and Batch Size to get the exact parameter count, parameter memory and total training memory for a fully-connected network.
  3. Learning Rate Schedule: enter an initial rate, pick constant / step / exponential / cosine decay, and set the decay rate, period and epoch count to see how the rate falls over training.
  4. Batch Size & Training: enter dataset size, batch size, epochs, GPU memory and your own measured per-step time and per-sample activation memory to get batches per epoch, total iterations, wall-clock estimate and the largest batch that fits.
  5. Regularization and Activation Functions: score a dropout / L2 / batch-norm / augmentation combination, or compare ReLU, Sigmoid, Tanh, GELU and Swish by formula, output range, gradient behaviour and relative compute cost.

What each input means

Analysis Type
Calculation mode to use.
Input Size
Number of input features, e.g. 784 for a flattened 28x28 image
Hidden Layers
Number of hidden layers. Must be a whole number of at least 1.
Hidden Units per Layer
Width of every hidden layer (this model uses a uniform width).
Output Size
Number of classes or regression targets
Batch Size
Samples per training step. Affects activation memory only, not the parameter count.
Model + Optimizer Memory (GB)
Weights, gradients and optimizer state. Roughly 16 bytes per parameter for fp32 + Adam; use the Network Architecture mode's Total Memory as a starting point.
Activation Memory per Sample (MB)
Peak activation memory for ONE training sample, measured from a batch-size-1 run. A small MLP is well under 1 MB; a 224x224 CNN is on the order of 10-100 MB.
Seconds per Batch (measured)
Wall-clock time for one training step, timed on your own hardware. There is no way to predict this from architecture alone -- run 20 steps and average.
Dropout Rate
0-1, typically 0.2-0.5

How this is calculated

Formula

Dense layer params = (inputs x outputs) + outputs | fp32 memory = 4 bytes x params

Worked example, using the default values

  1. Identify Input Parameters
    5 parameters
    Analysis Type = 0, Input Size = 784, Hidden Layers = 2, Hidden Units per Layer = 256, Output Size = 10 = 5 input(s) provided for Network Architecture mode
  2. Calculate Total Parameters
    Total Parameters
    269322 = 269322
  3. Calculate Parameter Memory
    Parameter Memory
    1.0273818969726562 = 1.0273818969726562
  4. Calculate Total Memory
    Total Memory
    4.268951416015625 = 4.268951416015625

Engine last updated . Checked against 6 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Which input changes Total Parameters the most?

It depends on the architecture, and the answer at the defaults surprises people: with 784 inputs and 256 hidden units, the input-to-first-hidden matrix alone is 200,704 of the 269,322 total parameters, so Input Size dominates. Hidden Units per Layer enters quadratically in each hidden-to-hidden matrix -- doubling it roughly quadruples those -- so it takes over once the network is deep enough, or the input narrow enough, for those matrices to outweigh the input projection.

Does Batch Size change how big my model is?

No -- Batch Size has zero effect on Total Parameters, since parameter count depends only on the network's architecture (layer sizes), not on how many samples you process at once. Batch Size does change Total Memory, because larger batches require more memory to hold activations during the forward and backward pass.

What do the other four Analysis Type modes calculate?

Learning Rate Schedule projects how your learning rate decays over training (constant, step, exponential, or cosine annealing) and recommends an optimizer; Batch Size & Training turns measured per-step timings and per-sample memory into batches per epoch, total training time and the largest batch your GPU can hold -- it asks you for those measurements rather than guessing them from the architecture; Regularization scores dropout, L2 weight decay, batch normalization, and data augmentation into an overall regularization strength and overfitting risk; and Activation Functions compares ReLU, Sigmoid, Tanh, GELU, and Swish by formula, gradient behavior, and relative compute cost.

What do the Complexity Rating bands mean?

Complexity Rating is a simple threshold classification of Total Parameters: Simple at 10,000 parameters or fewer, Medium above 10,000 up to 1 million, Large above 1 million up to 1 billion, and "Very Large (LLM scale)" above 1 billion parameters. It's a rough sizing label, not a claim that your specific architecture matches a real large language model's design.

Where do the Est. Training Time and Max Batch Size figures come from?

Entirely from what you enter. Training time is simply batches per epoch x epochs x your measured seconds-per-batch, and Max Batch Size is your GPU memory minus your model and optimizer memory, divided by your measured per-sample activation memory. The calculator does not try to predict step time or activation footprint from layer sizes -- those depend on the GPU, the precision, the kernels and the framework, and any figure derived from architecture alone would be wrong by orders of magnitude. Time twenty real training steps and measure peak memory at batch size 1, then enter those.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.