Evaluation of the accuracy, consistency, and scientific reliability of AI-powered Chatbots in endodontic practice

Acta Odontol Scand. 2026 Jun 4:85:305-310. doi: 10.2340/aos.v85.46148.

Abstract

Objective: This study aimed to evaluate the accuracy, consistency, and scientific reliability of two AI-powered chatbots-ChatGPT-3.5 and ChatGPT-4o-in clinical endodontic decision-making, using the recently published European Society of Endodontology (ESE) S3-level Clinical Practice Guideline as the gold standard reference.

Material and methods: Twenty-five dichotomous (yes/no) questions were developed based on the ESE guideline and presented to both chatbots across three time intervals, yielding 300 total responses. Each response was evaluated for accuracy and consistency, and the quality of the supporting references was assessed according to their journal ranking (Q1, Q2, others).

Results: Both ChatGPT versions demonstrated high internal consistency across repeated measurements. ChatGPT-3.5 showed 94.4% agreement (κ = 0.824; 95% confidence interval [CI]: 0.786-0.898; p < 0.001), whereas ChatGPT-4o demonstrated 98.9% agreement (κ = 0.937; 95% CI: 0.893-0.965; p < 0.001). The accuracy of ChatGPT-3.5 relative to the guideline-based answers was 81.4%, 88.9%, and 82.2% in the morning, afternoon, and evening sessions, respectively, while ChatGPT-4o achieved 82.9%, 83.3%, and 85.4%, respectively. No statistically significant differences were observed between the models across the three time intervals (p > 0.05). The proportion of Q1/Q2-ranked references was high and comparable between ChatGPT-3.5 (74-82%) and ChatGPT-4o (76-84%).

Conclusion: Both ChatGPT-3.5 and ChatGPT-4o demonstrated substantial alignment with the ESE S3-level clinical practice guideline. However, these findings should not be interpreted as definitive assessments of current clinical conversational AI systems, and further evaluation of evolving models is required.

MeSH terms

  • Endodontics*
  • Generative Artificial Intelligence*
  • Humans
  • Reproducibility of Results