Automating methodological quality assessment in orthopedic systematic reviews using large language models

J Orthop Surg (Hong Kong). 2026 May-Aug;34(2):10225536261459518. doi: 10.1177/10225536261459518. Epub 2026 Jun 7.

Abstract

BackgroundSystematic reviews represent the foundation of evidence-based orthopedic practice, yet their methodological rigor relies heavily on accurate and consistent methodological quality assessment. This step remains time-consuming, labor-intensive, and prone to subjectivity. Recent advances in large language models (LLMs) suggest potential for automating parts of evidence synthesis.PurposeThis study examined whether LLMs can perform AMSTAR-1-based methodological quality assessment evaluations in orthopedic systematic reviews with accuracy comparable to human experts.MethodsTen sports medicine knee reviews were analyzed using three LLMs-GPT-4o, GPT-5, and GPT Consensus-and their binary responses were compared against expert AMSTAR-1 ratings from a published umbrella review (110 decisions). An external validation set of four reviews published between 2022 and 2025 was included to assess generalizability and safeguard against information leakage.ResultsAgreement with human reviewers reached 87% for GPT-4o, 89% for GPT-5, and 90% for GPT Consensus; all models achieved 84% agreement in the validation set. Concordance was strongest for structured, explicitly reported domains such as a priori design, literature search, and study characteristics, and lowest for judgment-based items including grey literature inclusion, publication bias, and conflict of interest.ConclusionsLLMs cannot yet replace human reviewers, they can serve as reliable adjunct tools to enhance efficiency, transparency, and reproducibility in systematic review workflows within orthopedic research.

Keywords: artificial intelligence; evidence-based; large language model; sports.

MeSH terms

  • Humans
  • Large Language Models*
  • Orthopedics*
  • Systematic Reviews as Topic*