Skip to main content
Calcimator

Model Performance Metrics Calculator

Calculate accuracy, precision, recall, F1 score, MCC, AUC, and other classification performance metrics.

About this calculator

A binary classifier's predictions fall into four buckets against the true labels: true positives (TP, correctly flagged positive), true negatives (TN, correctly flagged negative), false positives (FP, wrongly flagged positive), and false negatives (FN, wrongly flagged negative) -- together, the confusion matrix. Every metric here is a different ratio built from those four counts, and each answers a different question. Precision (TP / (TP+FP)) asks "of everything the model flagged positive, how much was actually positive?" -- it's hurt only by false positives. Recall, also called sensitivity, (TP / (TP+FN)) asks "of everything that was actually positive, how much did the model catch?" -- it's hurt only by false negatives.

Those two pull in opposite directions: a model that flags everything positive gets perfect recall (it misses nothing) but terrible precision (most flags are wrong), and a model that only flags its most confident cases gets high precision but misses real positives, so accuracy alone -- (TP+TN)/total -- can be a misleading single number when the two classes are imbalanced, because a model that always predicts the majority class scores high accuracy while being useless at the task. F1 score (the harmonic mean of precision and recall) and Matthews Correlation Coefficient (MCC, which uses all four confusion-matrix cells including true negatives, and ranges from -1 to +1 with 0 meaning no better than random) are both attempts to compress the precision/recall tradeoff into one number that's harder to game by favoring one class. Which metric matters most is domain-specific: a spam filter cares more about precision (falsely blocking real email is worse than missing some spam), while a cancer-screening test cares more about recall (missing a real case is worse than a false alarm that gets ruled out later).

Inputs

Results

Accuracy

87.5%

Precision

89.47%

Recall (Sensitivity)

85%

F1 Score

87.18%

Matthews Correlation Coefficient75.09
AUC (Area Under Curve)87.5
How to Use This Calculator
  1. Enter the confusion matrix values: true positives, false positives, true negatives, false negatives.
  2. Review the derived metrics: accuracy, precision, recall, F1 score, and specificity.
  3. Check the Matthews Correlation Coefficient (MCC) for a single balanced summary that stays meaningful even when your classes are imbalanced.
  4. For imbalanced datasets, prioritize F1 score or AUC-ROC over raw accuracy.
  5. Compare metrics across models to select the best performer on your evaluation criteria.

How the result changes with False Negatives (FN)

False Negatives (FN)AccuracyPrecisionRecall (Sensitivity)
7.590.91%89.47%91.89%
1189.29%89.47%88.54%
2384.13%89.47%78.7%
3878.48%89.47%69.11%

What each input means

True Positives (TP)
Correctly predicted positive cases
True Negatives (TN)
Correctly predicted negative cases
False Positives (FP)
Incorrectly predicted as positive
False Negatives (FN)
Incorrectly predicted as negative

How this is calculated

Formula

F1 = 2 × (Precision × Recall) / (Precision + Recall)

Worked example, using the default values

  1. Identify Input Parameters
    4 parameters
    True Positives (TP) = 85, True Negatives (TN) = 90, False Positives (FP) = 10, False Negatives (FN) = 15 = 4 input(s) provided
  2. Calculate Accuracy
    Accuracy
    87.5 = 87.5%
  3. Calculate Precision
    Precision
    89.47 = 89.47%
  4. Calculate Recall
    Recall
    85 = 85%
  5. Calculate Matthews Correlation Coefficient
    Matthews Correlation Coefficient
    75.09 = 75.09
  6. Calculate AUC
    AUC
    87.5 = 87.5

Engine last updated . Checked against 2 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Why can accuracy be a misleading metric?

Accuracy weights every correct prediction equally regardless of which class it belongs to, so on an imbalanced dataset -- say 95% negative cases -- a model that predicts "negative" for everything scores 95% accuracy while never correctly catching a single positive case. Precision, recall, and MCC are all more informative in that situation because they expose exactly how the model is failing on the minority class, which accuracy alone hides.

Why do precision and recall move in opposite directions?

Because they're sensitive to different kinds of mistakes -- precision only gets hurt by false positives, and recall only gets hurt by false negatives. A model tuned to flag more cases as positive (to catch more true positives and raise recall) will typically also pick up more false positives along the way, which lowers precision, and vice versa for a more conservative model. This tradeoff is exactly what F1 score tries to balance into a single figure.

What does an MCC near 0 actually mean?

MCC ranges from -1 (total disagreement between predictions and actual labels) through 0 (no better than random guessing) to +1 (perfect prediction), and unlike accuracy or F1, it uses all four confusion-matrix cells including true negatives, which makes it harder to inflate by favoring one class. An MCC near 0 means the classifier's predictions carry essentially no real information about the true labels, even if accuracy looks high due to class imbalance.

Why does increasing true negatives raise accuracy but not precision or recall?

Accuracy counts every correct prediction, including true negatives, in its numerator and denominator, so correctly identifying more negative cases directly raises it. Precision and recall are both defined purely in terms of the positive class -- precision only looks at predicted positives, recall only at actual positives -- so neither formula includes true negatives at all, and adding more of them leaves both metrics completely unchanged.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Technology & Computing.