Performance of large language models in endodontics: accuracy, consistency, and benchmarking with consensus guidelines
BMC Oral Health, cilt.26, sa.1, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 26 Sayı: 1
- Basım Tarihi: 2026
- Doi Numarası: 10.1186/s12903-026-08269-8
- Dergi Adı: BMC Oral Health
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE, Directory of Open Access Journals, Natural Science Collection (ProQuest), Biological Science Database (ProQuest), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
- Anahtar Kelimeler: Artificial Intelligence, ChatGPT, Gemini, Large Language Models, Position Statements
- İstanbul Üniversitesi-Cerrahpaşa Adresli: Evet
Özet
Background: Large language model (LLM)-based chatbots are increasingly used in healthcare, yet their diagnostic accuracy, consistency, and temporal stability in endodontics remain insufficiently evaluated. This study aimed to assess and compare the performance of LLM-based chatbots using established international clinical guidelines. Methods: A diagnostic accuracy study with a repeated-measures design was conducted. Two LLMs (ChatGPT-5 and Gemini-2.5 Flash) were evaluated using 200 structured yes/no items derived from consensus-based international position statements covering the full scope of endodontic practice. Each model was tested weekly over three consecutive sessions, yielding 600 responses per model. Accuracy was assessed against reference answers, consistency was analyzed using Fleiss’ κ and Cohen’s κ, and logistic regression was performed to evaluate the effects of model type, week, and clinical domain. Results: ChatGPT demonstrated significantly higher overall accuracy than Gemini (92.8% vs. 84.8%; OR = 2.37; p = 0.004) and near-perfect reproducibility (κ = 0.95), whereas Gemini showed fair agreement (κ = 0.38). Both models performed well in structured clinical domains but showed reduced accuracy in areas related to restoration and pulpal disease management. Conclusions: LLM-based chatbots show potential as decision support and educational tools in endodontics. However, performance limitations in complex clinical domains highlight the continued need for expert oversight in clinical decision-making.