Multireader, multicase receiver operating characteristic analysis: an empirical comparison of five methods

Nancy A Obuchowski; Sergey V Beiden; Kevin S Berbaum; Stephen L Hillis; Hemant Ishwaran; Hae Hiang Song; Robert F Wagner

doi:10.1016/j.acra.2004.04.014

Multireader, multicase receiver operating characteristic analysis: an empirical comparison of five methods

Acad Radiol. 2004 Sep;11(9):980-95. doi: 10.1016/j.acra.2004.04.014.

Authors

Nancy A Obuchowski¹, Sergey V Beiden, Kevin S Berbaum, Stephen L Hillis, Hemant Ishwaran, Hae Hiang Song, Robert F Wagner

Affiliation

¹ Departments of Biostatistics and Epidemiology, Cleveland Clinic Foundation, Cleveland, OH 44195, USA. nobuchow@bio.ri.ccf.org

PMID: 15350579
DOI: 10.1016/j.acra.2004.04.014

Abstract

Rationale and objectives: Several statistical methods have been developed for analyzing multireader, multicase (MRMC) receiver operating characteristic (ROC) studies. The objective of this article is to increase awareness of these methods and determine if their results are concordant for published datasets.

Materials and methods: Data from three previously published studies were reanalyzed using five MRMC methods. For each method the 95% confidence intervals (CIs) for the mean of the readers' ROC areas for each diagnostic test, the P value for the comparison of the diagnostic tests' mean accuracies, and the 95% CIs for the mean difference in ROC areas of the diagnostic tests were reported.

Results: Important differences in P values and CIs were seen when using parametric versus nonparametric estimates of accuracy, and there were the expected differences for random-reader versus fixed-reader models. Controlling for these differences, the Dorfman-Berbaum-Metz (DBM), Obuchowski-Rockette, Beiden-Wagner-Campbell, and Song's multivariate Wilcoxon-Mann-Whitney (WMW) methods gave almost identical results for the fixed-reader model. For the random-reader model, the DBM, Obuchowski-Rockette, and Beiden-Wagner-Campbell methods yielded approximately the same inferences, but the CIs for the Beiden-Wagner-Campbell method tend to be broader. Ishwaran's hierarchical ROC sometimes yielded significance not found with other methods. Song's modification of DBM's jack-knifing algorithm sometimes led to different conclusions than the original DBM algorithm.

Conclusion: In choosing and applying MRMC methods, it is important to recognize: (1) the distinction between random-reader and fixed-reader models, the uncertainties accounted for by each, and thus the level of generalizeability expected from each; (2) assumptions made by the various MRMC methods; and (3) limitations of a five- or six-reader study when the reader variability is great.

Publication types

Comparative Study

MeSH terms

Analysis of Variance
Aortic Aneurysm / diagnostic imaging
Aortic Dissection / diagnostic imaging
Breast Neoplasms / diagnostic imaging
False Positive Reactions
Female
Humans
Lung Diseases, Interstitial / diagnostic imaging
Mammography
Models, Statistical
Multivariate Analysis
Observer Variation
ROC Curve*
Radiography*
Radiography, Thoracic
Signal Processing, Computer-Assisted
Statistics, Nonparametric
X-Ray Film