Assessing accuracy, readability & reliability of AI-generated patient leaflets on Descemet membrane endothelial keratoplasty

Eur J Ophthalmol. 2026 Jan;36(1):5-12. doi: 10.1177/11206721251367562. Epub 2025 Aug 28.

Abstract

PurposeThis study assessed the readability, reliability and accuracy of patient information leaflets on Descemet Membrane Endothelial Keratoplasty (DMEK), generated by seven large language models (LLMs). The aim was to determine which LLM produced the most patient-friendly, comprehensible and evidence-based leaflet, measured against a leaflet written by clinicians from a tertiary centre.MethodsEach LLM was given the prompt, "Make a patient information leaflet on Descemet Membrane Endothelial Keratoplasty (DMEK) surgery." Readability metrics (FKG, FRE, ARI, Gunning Fog), reliability metrics (DISCERN, PEMAT), misinformation detection and reference analysis were recorded for each response. A weighted scoring system normalised results on a 0-100% scale.ResultsThe clinician-generated leaflet scored the highest (92%). Claude 3.7 Sonnet had the top LLM score (77.8%), with strong readability and referencing. ChatGPT-4o followed closely (70.9%) but lacked references. Moderate scores for DeepSeek-V3, Perplexity AI and Google Gemini 2.0 Flash. ChatGPT-4 and Microsoft CoPilot scored the lowest due to limited reliability and misinformation.ConclusionsLLMs show promise in generating patient education material but vary in reliability and accuracy. Claude 3.7 Sonnet was the best performing LLM, though none matched in quality to the clinician-generated leaflet. LLM-generated leaflets therefore require clinician oversight before safe clinical use.

Keywords: ChatGPT; Claude; DISCERN; Large language models (LLMs); PEMAT; artificial intelligence in ophthalmology; descemet membrane endothelial keratoplasty; health communication; patient information leaflets (PILs); readability.

MeSH terms

  • Comprehension*
  • Descemet Stripping Endothelial Keratoplasty* / education
  • Humans
  • Pamphlets*
  • Patient Education as Topic* / methods
  • Patient Education as Topic* / standards
  • Reproducibility of Results