Analyses

Cohen’s kappa

Agreement between two raters, corrected for chance.

Edit on GitHub

When to use it

Use Cohen’s kappa when two raters assign the same cases to categories and you want agreement beyond what chance would produce.

Assumptions

Each row is one case rated twice. Both columns are categories with a small number of levels.

Running it in Tensr

Analyze → Descriptive Statistics → Crosstabs → Kappa. In chat: “Cohen’s kappa for rater A and rater B.”
Column A and column B are the two raters.

Options

Prop

Type

Reading the output

Two raters, 80 essays, categories 1, 2, and 3. Each rater keeps the true category most of the time and slips to a neighbour otherwise.

123
11340
212177
31818

Cohen's κ = 0.396 (fair agreement; κ < 0 poor, 0 ≤ κ < 0.2 slight, 0.2 ≤ κ < 0.4 fair, 0.4 ≤ κ < 0.6 moderate, 0.6 ≤ κ < 0.8 substantial, κ ≥ 0.8 almost perfect) for rater_a vs rater_b, n = 80. Cutoffs: Landis & Koch, 1977. The primary result is significant (p = < .001). This is large enough, in this sample, to treat the comparison this page is about as a real association rather than noise. Metrics: κ = 0.396; SE = 0.078; z = 5.079; 95% CI = [0.243, 0.549]; p-value = < .001; Observed agreement = 0.6.

Reporting (APA 7)

Cohen's κ = 0.396 (fair agreement; κ < 0 poor, 0 ≤ κ < 0.2 slight, 0.2 ≤ κ < 0.4 fair, 0.4 ≤ κ < 0.6 moderate, 0.6 ≤ κ < 0.8 substantial, κ ≥ 0.8 almost perfect) for rater_a vs rater_b, n = 80. Cutoffs: Landis & Koch, 1977. This result is significant (p = < .001). The effect is moderate, so report its size with the p-value.

Coming from SPSS

Crosstabs → Statistics → Kappa. SPSS prints kappa, an asymptotic standard error, and an approximate significance.