Skip to main content
Calcimator

Inter-Rater Reliability Calculator

Calculate Cohen's kappa coefficient to measure agreement between two raters beyond what would be expected by chance. Get an interpretation of your reliability score.

About this calculator

This calculator computes Cohen's Kappa (Cohen, 1960), a measure of rater agreement that corrects for how often two raters would agree purely by chance. Observed Agreement is Number of Agreements divided by Total Ratings, and Expected Agreement is the chance-agreement rate: rater1PositivePct × rater2PositivePct (both classify positive) plus their complements' product (both classify negative) (lines 13-16). At the calculator's default values, Rater 2 Positive Rate happens to sit at exactly 50%, and 0.5 × x + 0.5 × (1 − x) always equals 0.5 no matter what x is — so Expected Agreement, and therefore Kappa, is completely unaffected by Rater 1 Positive Rate at these specific defaults, even though the field visibly exists and matters at any other Rater 2 Positive Rate.

Kappa itself is (observed − expected) / (1 − expected) (line 21), and the engine labels the result using the standard Landis-Koch bands, from "Poor" below 0 up to "Almost Perfect" at 0.81 and above (lines 26-30). The calculator has no way to account for more than two raters or for rating categories beyond a simple positive/negative split.

Inputs

%
%

Results

Cohen's Kappa

0.6

InterpretationModerate
Observed Agreement80%
Expected Agreement50%
Agreements80
Disagreements20

Figures current as of 1960. Source: Cohen, J. (1960). "A Coefficient of Agreement for Nominal Scales." Educational and Psychological Measurement, 20(1), 37-46.

How to Use This Calculator
  1. Enter the number of raters and the number of items rated.
  2. Input the rating scale type (nominal, ordinal, interval).
  3. Enter the percentage of agreements observed across all rater pairs.
  4. Review the Cohen's Kappa or ICC coefficient along with its interpretation.
  5. If Kappa is below 0.60, schedule a rater calibration session and re-rate a subset of items.

How the result changes with Total Ratings

Total RatingsCohen's Kappa
502.2
751.13
1500.07
250-0.36

What each input means

Number of Agreements
The number of items where both raters assigned the same category or score.
Total Ratings
The total number of items rated by both raters.
Rater 1 Positive Rate (%)
Percentage of items that rater 1 classified as positive/yes. Used to calculate expected chance agreement.
Rater 2 Positive Rate (%)
Percentage of items that rater 2 classified as positive/yes. Used to calculate expected chance agreement.

How this is calculated

Worked example, using the default values

  1. Identify Input Parameters
    4 parameters
    Number of Agreements = 80, Total Ratings = 100, Rater 1 Positive Rate (%) = 55, Rater 2 Positive Rate (%) = 50 = 4 input(s) provided
  2. Calculate Cohen's Kappa
    Cohen's Kappa
    0.6 = 0.6
  3. Calculate Interpretation
    Interpretation
    Moderate = Moderate
  4. Calculate Observed Agreement
    Observed Agreement
    80 = 80

Figures and sources

Engine last updated . Checked against 2 independently-derived tests — how we verify calculators. Built by Paul Gunder, a software engineer, not a licensed financial, medical, or legal professional.

Frequently Asked Questions

Does Rater 1 Positive Rate affect Kappa?

Not at the calculator's default values, though it isn't inert in general. Rater 2 Positive Rate defaults to exactly 50%, and 0.5 × x + 0.5 × (1 − x) always equals 0.5 regardless of x (lines 14-16), so Expected Agreement — and therefore Kappa — doesn't move when Rater 1 Positive Rate changes only because Rater 2's rate happens to sit at that exact midpoint.

Would Rater 1 Positive Rate matter if I changed Rater 2 Positive Rate away from 50%?

Yes. Expected Agreement is rater1PositivePct × rater2PositivePct plus their complements' product (lines 14-16); that expression only collapses to a constant when rater2PositivePct is exactly 0.5. At any other Rater 2 Positive Rate, Rater 1 Positive Rate does change Expected Agreement and, through it, Kappa.

What does a Kappa of 0.6 mean in plain terms?

The engine labels it "Moderate" agreement (line 28, kappa ≥ 0.41) — meaningfully better than chance-level agreement, but short of the "Substantial" (≥0.61) or "Almost Perfect" (≥0.81) bands used for higher scores under the standard Landis-Koch interpretation scale. Kappa itself comes from Jacob Cohen's 1960 paper "A Coefficient of Agreement for Nominal Scales" (Educational and Psychological Measurement), which defined it specifically to correct raw percent agreement for the agreement two raters would rack up by chance alone.

How many raters and rating categories does this calculator support?

Exactly two raters and a simple positive/negative classification — the formulas for Observed Agreement, Expected Agreement, and Kappa (lines 10-21) are all built around that two-rater, two-category structure and don't generalize to comparisons among three or more raters or to ordinal/multi-category rating scales.

The questions that sit next to this one — chosen by subject, including calculators filed under a different category.

More in Math & Statistics.