Overfitting Calculator
Calculate overfitting metrics, generalization gap, bias-variance trade-off, and model complexity analysis.
About this calculator
This calculator computes the Overfitting Gap as simply Training Accuracy minus Validation Accuracy -- the standard, most direct signal that a model has memorized training examples rather than learned generalizable patterns. A larger gap means the model performs much better on data it has already seen than on new data, the hallmark of overfitting. The calculator reuses this same figure as the Generalization Gap and as Variance in a simplified bias-variance framing, while Bias is estimated as 100 minus Training Accuracy (a proxy for how much the model fails to fit even its own training data -- high bias means underfitting).
Complexity Ratio (Model Complexity in thousands of parameters, divided by Dataset Size) is a rough capacity-to-data measure: a rule of thumb in machine learning is that models need roughly 10 training samples per parameter to generalize well, which is also how the calculator derives Optimal Complexity (Dataset Size divided by 10). Effective Regularization multiplies Regularization Strength by Model Complexity, and Expected Validation Accuracy is a simplified projection of what regularization could recover, assuming it closes a fraction of the gap proportional to the regularization coefficient -- Regularization Strength does not change the reported Overfitting Gap itself, since that figure reflects the accuracies as entered, only the separate Expected Validation Accuracy projection. These are simplified proxies, not a trained model's actual measured bias and variance decomposition, which in practice require repeated training runs across resampled datasets.
Inputs
Results
Overfitting Gap
10%
Generalization Gap
10%
How to Use This Calculator
- Enter Training Accuracy and Validation Accuracy (%) from your model's evaluation.
- Enter Model Complexity (thousand parameters) and Dataset Size (training samples).
- Set Regularization Strength to see its projected effect on Expected Validation Accuracy.
- Review the Overfitting Gap and Generalization Gap — a large positive value signals overfitting.
- Compare Complexity Ratio against Optimal Complexity to judge whether the model has enough training data for its size.
How the result changes with Training Accuracy
| Training Accuracy | Overfitting Gap | Generalization Gap |
|---|---|---|
| 48% | -37% | -37% |
| 71% | -14% | -14% |
| 143% | 58% | 58% |
| 238% | 153% | 153% |
What each input means
- Training Accuracy
- Model accuracy on training set
- Validation Accuracy
- Model accuracy on validation set
- Model Complexity
- Number of model parameters
- Dataset Size
- Number of training samples
- Regularization Strength
- Regularization coefficient
How this is calculated
Formula
Overfitting Gap = Training Accuracy - Validation AccuracyWorked example, using the default values
- Identify Input Parameters5 parametersTraining Accuracy = 95, Validation Accuracy = 85, Model Complexity = 50, Dataset Size = 10000, Regularization Strength = 0.01 = 5 input(s) provided
- Calculate Overfitting GapOverfitting Gap10 = 10%
- Calculate Generalization GapGeneralization Gap10 = 10%
- Calculate VarianceVariance10 = 10
- Calculate BiasBias5 = 5
Engine last updated . Checked against 2 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.
Frequently Asked Questions
How is the Overfitting Gap calculated?
Overfitting Gap is simply Training Accuracy minus Validation Accuracy. A larger positive gap means the model scores much higher on data it was trained on than on unseen validation data, which is the classic symptom of overfitting -- the model has memorized specifics of the training set rather than learning patterns that generalize.
Does increasing Regularization Strength change the reported Overfitting Gap?
No. The Overfitting Gap is computed directly from the Training Accuracy and Validation Accuracy you enter, with no dependence on Regularization Strength. Regularization Strength only feeds the separate Expected Validation Accuracy output, a projection of what accuracy regularization might recover, not a recalculation of the gap you actually measured.
Where does the 'Optimal Complexity' figure come from?
It applies a common machine-learning rule of thumb -- roughly 10 training samples needed per model parameter for reliable generalization -- by dividing Dataset Size by 10. This is a rough capacity guideline, not a guarantee: the right amount of data per parameter varies substantially by model architecture, task difficulty, and regularization approach.
What is the difference between Bias and Variance in this calculator?
Variance here is set equal to the Overfitting Gap (high gap = high variance, meaning the model's predictions vary too much between the training set and new data). Bias is estimated as 100 minus Training Accuracy, a proxy for how much the model underfits even its own training data. These are simplified single-number estimates, not a full bias- variance decomposition from repeated resampled training.
How does Model Complexity affect the Complexity Ratio?
Complexity Ratio is Model Complexity (in thousands of parameters) divided by Dataset Size, so raising Model Complexity while holding Dataset Size fixed always raises the ratio proportionally. A higher ratio signals a model with more parameters relative to available training examples, which raises overfitting risk per the same 10-samples-per- parameter guideline used for Optimal Complexity.
Related Calculators
The questions that sit next to this one — chosen by subject, including calculators filed under a different category.
Cross-Validation Calculator
Calculate k-fold cross-validation splits, train-test splits, and data utilization for machine learning.
Machine Learning & AIModel Performance Metrics Calculator
Calculate accuracy, precision, recall, F1 score, MCC, AUC, and other classification performance metrics.
Machine Learning & AIGradient Descent Calculator
Calculate gradient descent parameters, convergence rate, effective learning rate, and training time estimates.
MLOps & AI CostingFeature Importance Calculator
Estimate how many features to keep, overfitting risk, and expected variance retention based on dataset size, model type, and correlation threshold.
More in Technology & Computing.