Kendall’s W
How much a set of rankings agree with each other.
When to use it
Use Kendall’s W when several judges rank the same items, or several repeated scores should move together. Friedman’s test is the related hypothesis test. W is the effect size, from 0 (no agreement) to 1 (complete agreement).
Assumptions
Each column is one ranking or one repeated measure. Rows are the items or the people, depending on how you laid the data out. On this procedure Tensr treats columns as judges and rows as items.
Running it in Tensr
Options
Prop
Type
Reading the output
Four judges scoring six essays, with almost no shared ranking. Concordance was built to be weak.
Kendall's W = 0.286 (small concordance) for judges judge1, judge2, judge3, judge4 across 6 subjects. The primary result is not significant (p = .335). This is not large enough to treat the comparison this page is about as a reliable association. The result is non-significant: the data are still compatible with no effect. Metrics: W = 0.286; χ² = 5.714; df = 5; p-value = .335.
Reporting (APA 7)
Kendall's W = 0.286 (small concordance) for judges judge1, judge2, judge3, judge4 across 6 subjects. This result is not significant (p = .335). Report the estimate with that p, and do not describe the pattern as a reliable effect.
Coming from SPSS
Analyze → Scale → Kendall’s W. Friedman’s ANOVA is the related nonparametric test when the rows are people and the columns are repeated measures.