Neural Network Parameters Calculator
Calculate total parameters, weights, biases, and memory requirements for neural network architectures.
About this calculator
This calculator sums the weight matrices of a fully-connected (dense) feedforward network: Input Neurons times Neurons per Hidden Layer for the first connection, Neurons per Hidden Layer squared for each additional hidden-to-hidden connection, and Neurons per Hidden Layer times Output Neurons for the final layer, then adds one bias term per neuron when Bias Terms is set to "With Bias." Setting Hidden Layers to 0 removes both of those hidden-layer weight blocks entirely and connects Input Neurons directly to Output Neurons with a single weight matrix, modeling a network with no hidden layers at all rather than reusing the one-hidden-layer arithmetic. Because the hidden-to-hidden term scales with the square of Neurons per Hidden Layer while every other term scales linearly with its neuron counts, Neurons per Hidden Layer is the input with the most leverage over Total Weights for a typical multi-hidden-layer network -- doubling it can more than double the parameter count once two or more hidden layers are in play. Hidden Layers adds a full square block of weights (Neurons per Hidden Layer squared) for every additional hidden-to-hidden connection, so parameter count grows roughly linearly with layer count once the layer size is fixed.
Memory Required simply multiplies Total Parameters by 4 bytes, assuming 32-bit floating point storage; it does not account for optimizer state (which can multiply working memory by 2-3x during training for algorithms like Adam), activation memory, or lower-precision formats like float16 or int8 quantization, all of which change real deployed or in-training memory substantially from this raw parameter-count figure. The formula also assumes a plain dense architecture -- convolutional, attention, or embedding layers use entirely different parameter-counting rules and are not represented here.
Inputs
Results
Total Parameters
118,282
Memory Required
0.45 MB
How to Use This Calculator
- Enter the network architecture: number of layers and units per layer.
- Input the input feature dimension and output classes or units.
- Review the total parameter count (weights + biases) for each layer and the network total.
- Compare parameter count to your dataset size: aim for at least 10 samples per parameter to avoid overfitting.
- Use the parameter count to estimate memory footprint and training compute requirements.
How the result changes with Neurons per Hidden Layer
| Neurons per Hidden Layer | Total Parameters | Memory Required |
|---|---|---|
| 64 | 55,050 | 0.21 MB |
| 96 | 85,642 | 0.33 MB |
| 192 | 189,706 | 0.72 MB |
| 320 | 357,130 | 1.36 MB |
What each input means
- Input Neurons
- Number of input neurons
- Hidden Layers
- Number of hidden layers
- Neurons per Hidden Layer
- Number of neurons in each hidden layer
- Output Neurons
- Number of output neurons
- Bias Terms
- Whether to include bias terms
How this is calculated
Formula
Parameters = (Input × Hidden) + (Hidden × Hidden) + (Hidden × Output) + BiasesWorked example, using the default values
- Identify Input Parameters5 parametersInput Neurons = 784, Hidden Layers = 2, Neurons per Hidden Layer = 128, Output Neurons = 10, Bias Terms = 1 = 5 input(s) provided
- Calculate Total ParametersTotal Parameters118282 = 118282
- Calculate Memory RequiredMemory Required0.45 = 0.45
- Calculate Total WeightsTotal Weights118016 = 118016
- Calculate Total BiasesTotal Biases266 = 266
Engine last updated . Checked against 3 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.
Frequently Asked Questions
Why does Neurons per Hidden Layer affect Total Parameters more than the other inputs?
Every pair of adjacent hidden layers contributes a weight block sized Neurons per Hidden Layer squared, while Input Neurons and Output Neurons only multiply linearly against a single layer's width. With two or more hidden layers, that squared term typically dominates the total, so widening the hidden layers moves the parameter count faster than widening the input or output layer by the same factor.
Does Memory Required account for training memory, not just the model file size?
No. It only multiplies Total Parameters by 4 bytes per float32 value, which approximates the raw model weights on disk or in inference memory. Training typically needs several times more memory once optimizer state (like Adam's per-parameter moving averages), gradients, and activations for backpropagation are included, so treat this figure as a lower bound during training.
How much do bias terms add to the total parameter count?
Bias terms add one value per neuron in every hidden and output layer -- typically a small fraction of the total compared with the weight matrices, since a weight matrix connecting two layers has far more entries than either layer has neurons. Turning Bias Terms off removes that small addition entirely but usually changes Total Parameters only modestly for realistically sized layers.
Does this formula work for convolutional or transformer architectures?
No. It only models a fully-connected dense network where every neuron in one layer connects to every neuron in the next. Convolutional layers share weights across spatial positions (far fewer parameters per layer than a dense equivalent), and transformer attention layers have their own parameter-counting rules based on embedding dimension and head count, so neither is represented by this calculator's arithmetic.
Related Calculators
The questions that sit next to this one — chosen by subject, including calculators filed under a different category.
Gradient Descent Calculator
Calculate gradient descent parameters, convergence rate, effective learning rate, and training time estimates.
Machine Learning & AIModel Performance Metrics Calculator
Calculate accuracy, precision, recall, F1 score, MCC, AUC, and other classification performance metrics.
Machine Learning & AINeural Network Trainer Calculator
Complete neural network training analysis. Architecture design, learning rates, batch sizes, regularization, and activation functions.
More in Technology & Computing.