Skip to main content
Calcimator

Model Performance Calculator

Complete ML model evaluation. Classification metrics, regression analysis, cross-validation, ROC/AUC curves, and model comparison.

About this calculator

Model performance evaluation splits into fundamentally different questions depending on what kind of model you're checking, which is why this calculator offers five separate evaluation modes rather than one universal score. Classification Metrics starts from a confusion matrix (true/false positives and negatives) and derives accuracy, precision, recall, and F1 score -- the standard toolkit for a model that predicts discrete categories. Accuracy alone can be misleading on imbalanced data (a model that always predicts "no fraud" can be 99% accurate on a dataset where fraud is rare), which is why precision (of the cases flagged positive, how many really were) and recall (of the actual positives, how many were caught) are reported separately, and F1 combines them via their harmonic mean so a model can't hide a weak recall behind strong precision or vice versa. Regression Metrics instead evaluates continuous predictions using RMSE, MAE, and R² -- how much of the actual outcome's variance the model explains -- computed here from a mean squared error and variance you provide directly, with MAE estimated from RMSE using the exact relationship that holds when prediction errors are normally distributed (MAE = RMSE x sqrt(2/pi)), an approximation, not a measurement of your actual error distribution.

Cross-Validation Analysis reports the stability of a model's score across folds and a 95% confidence interval on the mean. ROC/AUC Analysis computes true/false positive rates using a smooth parametric curve shaped to match your entered AUC value -- illustrative of the shape a given AUC typically produces, not your model's actual measured ROC curve. Model Comparison runs an approximate two- proportion statistical test (a normal-approximation p-value) to judge whether two models' accuracy difference is likely real or noise, alongside an efficiency comparison accounting for each model's parameter count; this test treats the two accuracies as independent samples, which typically understates significance when both models were actually evaluated on the SAME test set (the common case) -- two classifiers usually agree on "easy" examples and disagree mainly on "hard" ones, a positive correlation that a proper paired test exploits and an independent-samples test ignores. A true paired test for that situation needs the count of examples where the two models disagreed, which this calculator's summary-accuracy inputs don't capture.

Progress0%

Step 1 of 2

How to Use This Calculator
  1. Select an Evaluation Type: Classification Metrics, Regression Metrics, Cross-Validation, ROC/AUC Analysis, or Model Comparison.
  2. For Classification Metrics, enter your confusion-matrix counts (True/False Positives and Negatives) from comparing predictions to true labels on a test set.
  3. For Regression Metrics, enter the mean squared error, variance, and mean actual/predicted values from your model's evaluation.
  4. For Cross-Validation, ROC/AUC, or Model Comparison, enter the corresponding summary statistics (fold scores, AUC, or the two models' accuracies and parameter counts).
  5. Review the mode's outputs and interpretation label -- each mode's explainer and FAQ note what its numbers do and don't measure (e.g. the ROC curve is illustrative of the given AUC's shape, not your model's actual measured curve).

What each input means

Evaluation Type
Calculation mode to use.
Number of Samples
Must comfortably exceed the 10 assumed features (Adjusted R² is undefined once samples fall to features + 1).
AUC Score
Area Under ROC Curve
Positive Class Rate
Prevalence of positive class

How this is calculated

Formula

F1 = 2×(P×R)/(P+R) | R² = 1 - SS_res/SS_tot | AUC = ∫TPR d(FPR)

Worked example, using the default values

  1. Identify Input Parameters
    4 parameters
    True Positives (TP) = 85, True Negatives (TN) = 90, False Positives (FP) = 10, False Negatives (FN) = 15 = 4 input(s) provided
  2. Calculate Accuracy
    Accuracy = (TP + TN) / Total
    (85 + 90) / 200 = 87.5%
  3. Calculate F1 Score
    F1 = 2 × (Precision × Recall) / (Precision + Recall)
    2 × (0.8947 × 0.85) / (0.8947 + 0.85) = 87.2%
  4. Calculate Precision
    Precision = TP / (TP + FP)
    85 / (85 + 10) = 89.5%
  5. Calculate Recall
    Recall = TP / (TP + FN)
    85 / (85 + 15) = 85.0%

Engine last updated . Checked against 5 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Why does this calculator report both precision and recall instead of just accuracy?

Because accuracy alone can be badly misleading, especially when the classes are imbalanced. A spam filter that never flags anything as spam is highly "accurate" on a dataset where most email isn't spam, but useless at its actual job. Precision tells you how trustworthy a positive prediction is (of everything flagged positive, how much really was), while recall tells you how complete the model's positive predictions are (of everything actually positive, how much got caught) -- reporting them separately exposes tradeoffs that a single accuracy number hides entirely.

Is the MAE shown in Regression Metrics mode a real measurement?

No -- it's estimated from RMSE using the relationship MAE = RMSE x sqrt(2/pi), which holds exactly only when prediction errors are normally distributed. This calculator doesn't take individual prediction errors as input, only a summary MSE and variance, so it can't measure MAE directly; it derives an estimate under that normality assumption instead. If your model's actual errors are skewed or heavy-tailed rather than normally distributed, the real MAE could differ meaningfully from this estimate.

Does the ROC/AUC mode plot my model's actual ROC curve?

No -- it plots a smooth, parametric curve shaped to match the AUC value you enter, using a simplified mathematical model rather than your classifier's real threshold-by-threshold true/false positive rates. It's useful for visualizing roughly what shape a given AUC typically corresponds to and for exploring how the optimal threshold and lift change with AUC, but it is not a substitute for computing an actual ROC curve from your model's real predictions and true labels.

How does Model Comparison decide if one model is really better than another?

It runs an approximate two-proportion statistical significance test on the accuracy difference between the two models, using a normal-approximation p-value based on the pooled accuracy and test set size, then checks it against the conventional p < 0.05 threshold. This test assumes the two accuracies are independent samples; if you actually evaluated both models on the SAME test set (the usual setup), the correct test is a paired one (like McNemar's test, which needs the count of examples the two models disagreed on) and would typically be more sensitive than this approximation, not less -- so a "not significant" result here can be a false negative for a genuinely different, same-test-set comparison. It also separately reports an efficiency comparison -- accuracy relative to parameter count -- because a model that's not statistically significantly more accurate but is far smaller may still be the better practical choice, which is why the recommendation logic weighs both significance and efficiency, not accuracy alone.

Why do the classification metrics use True Positives, False Positives, etc. instead of asking for my raw predictions?

Because a confusion matrix -- counts of true positives, true negatives, false positives, and false negatives -- is a complete, compact summary of a binary classifier's performance at one decision threshold; every metric in this mode (accuracy, precision, recall, F1, MCC, specificity) is derivable from those four counts alone. You get those counts by comparing your model's predictions to the true labels on a test set and counting how each prediction falls into the four categories, which most ML libraries compute for you directly.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.