Kappa coefficients in medical research

Helena Chmura Kraemer; Vyjeyanthi S Periyakoil; Art Noda

doi:10.1002/sim.1180

Kappa coefficients in medical research

Stat Med. 2002 Jul 30;21(14):2109-29. doi: 10.1002/sim.1180.

Authors

Helena Chmura Kraemer¹, Vyjeyanthi S Periyakoil, Art Noda

Affiliation

¹ Department of Psychiatry and Behavioral Sciences, Stanford University School of Medicine, Stanford, California, U.S.A. hck@leland.stanford.edu

PMID: 12111890
DOI: 10.1002/sim.1180

Abstract

Kappa coefficients are measures of correlation between categorical variables often used as reliability or validity coefficients. We recapitulate development and definitions of the K (categories) by M (ratings) kappas (K x M), discuss what they are well- or ill-designed to do, and summarize where kappas now stand with regard to their application in medical research. The 2 x M(M>/=2) intraclass kappa seems the ideal measure of binary reliability; a 2 x 2 weighted kappa is an excellent choice, though not a unique one, as a validity measure. For both the intraclass and weighted kappas, we address continuing problems with kappas. There are serious problems with using the K x M intraclass (K>2) or the various K x M weighted kappas for K>2 or M>2 in any context, either because they convey incomplete and possibly misleading information, or because other approaches are preferable to their use. We illustrate the use of the recommended kappas with applications in medical research.

Kappa coefficients in medical research

Authors

Affiliation

Abstract

Publication types

MeSH terms

Grants and funding