Cohen’s kappa
Agreement between two raters, corrected for chance.
When to use it
Use Cohen’s kappa when two raters assign the same cases to categories and you want agreement beyond what chance would produce.
Assumptions
Each row is one case rated twice. Both columns are categories with a small number of levels.
Running it in Tensr
Options
Prop
Type
Reading the output
Two raters, 80 essays, categories 1, 2, and 3. Each rater keeps the true category most of the time and slips to a neighbour otherwise.
| 1 | 2 | 3 | |
|---|---|---|---|
| 1 | 13 | 4 | 0 |
| 2 | 12 | 17 | 7 |
| 3 | 1 | 8 | 18 |
Cohen's κ = 0.396 (fair agreement; κ < 0 poor, 0 ≤ κ < 0.2 slight, 0.2 ≤ κ < 0.4 fair, 0.4 ≤ κ < 0.6 moderate, 0.6 ≤ κ < 0.8 substantial, κ ≥ 0.8 almost perfect) for rater_a vs rater_b, n = 80. Cutoffs: Landis & Koch, 1977. The primary result is significant (p = < .001). This is large enough, in this sample, to treat the comparison this page is about as a real association rather than noise. Metrics: κ = 0.396; SE = 0.078; z = 5.079; 95% CI = [0.243, 0.549]; p-value = < .001; Observed agreement = 0.6.
Reporting (APA 7)
Cohen's κ = 0.396 (fair agreement; κ < 0 poor, 0 ≤ κ < 0.2 slight, 0.2 ≤ κ < 0.4 fair, 0.4 ≤ κ < 0.6 moderate, 0.6 ≤ κ < 0.8 substantial, κ ≥ 0.8 almost perfect) for rater_a vs rater_b, n = 80. Cutoffs: Landis & Koch, 1977. This result is significant (p = < .001). The effect is moderate, so report its size with the p-value.
Coming from SPSS
Crosstabs → Statistics → Kappa. SPSS prints kappa, an asymptotic standard error, and an approximate significance.