Large language models in methodological quality evaluation of radiomics research based on METRICS: ChatGPT vs NotebookLM vs radiologist

Eur J Radiol. 2025 Mar:184:111960. doi: 10.1016/j.ejrad.2025.111960. Epub 2025 Jan 29.

Abstract

Objectives: This study aimed to evaluate the effectiveness of large language models (LLM) in assessing the methodological quality of radiomics research, using METhodological RadiomICs Score (METRICS) tool.

Methods: This study included open access radiomic research articles published in 2024 across various journals and a preprint repository, all under the Creative Commons Attribution License. Each study was independently evaluated using METRICS by two LLMs, ChatGPT-4 and NotebookLM, and a consensus assessment performed by two radiologists with expertise in radiomics research.

Results: A total of 48 open access articles were included in this study. ChatGPT-4, NotebookLM, and human readers achieved median scores of 79.5 %, 61.6 %, and 69.0 %, respectively, with a statistically significant difference across these evaluations (p < 0.05). Pairwise comparisons indicated no statistically significant difference for NotebookLM vs human experts (p > 0.05), in contrast to other pairs (p < 0.05). Intraclass correlation coefficient (ICC) for ChatGPT-4 and human experts was 0.563 (95 % CI: 0.050---0.795), corresponding to poor to good agreement. The ICC for ChatGPT-4 and NotebookLM and for human experts and NotebookLM were 0.391 (95 % CI: -0.031---0.665) and 0.555 (95 % CI: 0.326---0.723), respectively, indicating poor to moderate agreement. LLMs completed the tasks in a significantly shorter time (p < 0.05). In item-wise reliability analysis, ChatGPT-4 generally demonstrated higher consistency than NotebookLM.

Conclusion: LLMs hold promise for automatically evaluating the quality of radiomics research using METRICS, a new tool that is relatively more complex yet comprehensive compared to its counterparts. However, substantial improvements are needed for full alignment with human experts.

Keywords: Artificial intelligence; Large language models; Machine learning; Radiomics; Texture analysis.

Publication types

  • Comparative Study

MeSH terms

  • Biomedical Research*
  • Generative Artificial Intelligence
  • Humans
  • Large Language Models
  • Natural Language Processing*
  • Radiologists
  • Radiology*
  • Radiomics
  • Reproducibility of Results