Accuracy and Consistency of Three Large Language Models on Fixed Prosthodontics Questions
Consistency of Large Language Models on Fixed Prosthodontics Questions
DOI:
https://doi.org/10.54393/pbmj.v9i8.1434Keywords:
Artificial Intelligence, Large Language Models, Prosthodontics, Fixed ProsthodonticsAbstract
Large language models (LLMs) are increasingly used in dental education and clinical settings, but their accuracy and consistency in fixed prosthodontics remain uncertain. Objectives: To compare the accuracy (using a strict three-attempt criterion) and repeated response consistency of ChatGPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.8 in answering binary fixed prosthodontics questions across six clinical domains. Methods: Forty binary questions covering six core fixed prosthodontic domains were developed by the principal investigator. Two experienced faculty members independently answered the questions, and disagreements were resolved through discussion to establish the gold standard. Each question was submitted thrice to each LLM in a new chat session on the same day. Accuracy was the percentage of correct responses compared with the gold standard across all three attempts, while consistency was the percentage of identical responses across three attempts, regardless of correctness. Accuracy and consistency were expressed as percentages with 95% confidence intervals. Cochran’s Q test, Cohen’s kappa, and Fleiss’ kappa were used for statistical analysis. Results: Pre-consensus examiner agreement was moderate, 77.5% (Cohen’s κ=0.55). Claude Opus 4.8 achieved the highest accuracy (95% and consistency (97.5%, followed by ChatGPT-5.4 87.5% accuracy and 92.5% consistency and Gemini 3.1 Pro (85% accuracy and 87.5% consistency. No significant difference was observed in accuracy among models (Q=5.20, p=0.074). Intra-model consistency for three models was almost perfect (κ=0.83–0.97). Conclusions: The evaluated LLMs demonstrated high accuracy and consistency for structured fixed prosthodontics questions, with Claude Opus 4.8 highest, though not significant.
References
Al Hendi KD, Alyami MH, Alkahtany M, Dwivedi A, Alsaqour HG. Artificial Intelligence in Prosthodontics. Bioinformation. 2024 Mar; 20(3): 238. doi: 10.6026/973206300200238. DOI: https://doi.org/10.6026/973206300200238
Alfaraj A, Limones Á, Ahmad S, Aljubairah F, Albalaw S, Albesher M et al. Harnessing AI in Prosthodontics and Implant Dentistry: An Umbrella Review of Systematic Evidence. Journal of Prosthodontics. 2026 Feb; 35(2): 127-42. doi: 10.1111/jopr.70091. DOI: https://doi.org/10.1111/jopr.70091
Roustan D and Bastardot F. The Clinicians’ Guide to Large Language Models: A General Perspective with a Focus on Hallucinations. Interactive Journal of Medical Research. 2025 Jan 28; 14(1): e59823. doi: 10.2196/59823. DOI: https://doi.org/10.2196/59823
Kong HJ and Kim YL. Application of Artificial Intelligence in Dental Crown Prosthesis: A Scoping Review. BioMed Central Oral Health. 2024 Aug; 24(1): 937. doi: 10.1186/s12903-024-04657-0. DOI: https://doi.org/10.1186/s12903-024-04657-0
Çakar M, Avcı AT, Düzgün S, Aslan T, Hekimoğlu KN. Assessment of the Accuracy of Modern Artificial Intelligence Chatbots in Responding to Endodontic Queries. Australian Endodontic Journal. 2025 Dec; 51(3): 732-9. doi: 10.1111/aej.70012. DOI: https://doi.org/10.1111/aej.70012
Lafourcade C, Kerouredan O, Ballester B, Richert R. Accuracy, Consistency, and Contextual Understanding of Large Language Models in Restorative Dentistry and Endodontics. Journal of Dentistry. 2025 Jun; 157: 105764. doi: 10.1016/j.jdent.2025.105764. DOI: https://doi.org/10.1016/j.jdent.2025.105764
Kuşçu HY and Görüş Z. Performance of Large Language Models on Prosthodontics Questions of the Dentistry Specialization Examination: A Comparative Analysis (2014–2024). Anatolian Current Medical Journal; 7(6):893-9. doi: 10.38053/acmj.1789931. DOI: https://doi.org/10.38053/acmj.1789931
Sozen Yanik I, Sahin Hazir D, Bilgin Avsar D. Cross-Lingual Performance of Large Language Models in Maxillofacial Prosthodontics: A Comparative Evaluation. BioMed Central Oral Health. 2025 Oct; 25(1): 1630. doi: 10.1186/s12903-025-07035-6. DOI: https://doi.org/10.1186/s12903-025-07035-6
Pulkundwar P, Dhanawade V, Yadav R, Sonkar M, Asurlekar M, Rathod S. A Concise Review of Hallucinations in LLMs and their Mitigation. 2025 Dec.
Madfa AA, Alshammari AF, Anazi BA, Alenezi YE, Alkurdi KA. Accuracy and Reliability of Manus, ChatGPT, and Claude in Case-Based Dental Diagnosis. Frontiers in Oral Health. 2025; 6: 1686090. doi: 10.3389/froh.2025.1686090. DOI: https://doi.org/10.3389/froh.2025.1686090
Rosenstiel SF. Contemporary Fixed Prosthodontics 6E, South Asia Edition-E-Book: Contemporary Fixed Prosthodontics 6E, South Asia Edition-E-Book. Elsevier Health Sciences. 2022 Nov.
Geduk G, Hasırcı UC, Kusay DD, Aras RÇ, Çapar İ, Altın E et al. A Comparative Analysis of the Performance of Large Language Models in the Dentistry Specialty Examination. Scientific Reports. 2026 Jan; 16(1): 6739. doi: 10.1038/s41598-026-37800-8. DOI: https://doi.org/10.1038/s41598-026-37800-8
Erdal SG, Güdül AB, Köroğlu A. The Effectiveness of Large Language Models in Dental Specialty Questions: A Comparative Study in the Field of Prosthodontics. BioMed Central Medical Education. 2026 Feb; 26(1): 455. doi: 10.1186/s12909-026-08808-5. DOI: https://doi.org/10.1186/s12909-026-08808-5
Sears LM, Koseoglu M, Antonopoulou S, Lee SK, Kase M, Colebeck A et al. Assessing the Ability of Large Language Models to Summarize and Generate Maxillofacial Prosthetic Treatment Options. Journal of Prosthodontics. 2026 Apr. doi: 10.1111/jopr.70136. DOI: https://doi.org/10.1111/jopr.70136
Gheisarifar M, Shembesh M, Koseoglu M, Fang Q, Afshari FS, Yuan JC et al. Evaluating the Validity and Consistency of Artificial Intelligence Chatbots in Responding to Patients’ Frequently Asked Questions in Prosthodontics. The Journal of Prosthetic Dentistry. 2025 Jul; 134(1): 199-206. doi: 10.1016/j.prosdent.2025.03.009. DOI: https://doi.org/10.1016/j.prosdent.2025.03.009
Koyuncuoglu CZ, Selcuker AH, Ozyilmaz E. The Reliability of Answers from Four Different AI Chatbots on Periodontology Theoretical Exam Questions: An Evaluation in Dental Education. BioMed Central Oral Health. 2025 Dec; 26(1): 114. doi: 10.1186/s12903-025-07387-z. DOI: https://doi.org/10.1186/s12903-025-07387-z
Chutia J, Rupali R, Jain M, Angel A, Agarwal P, Sawhney C. Faculty Perspectives on Integrating Artificial Intelligence into Orthodontic and Interdisciplinary Dental Education: Opportunities, Challenges, And Strategies. Cureus. 2025 Sep; 17(9). doi: 10.7759/cureus.92877. DOI: https://doi.org/10.7759/cureus.92877
Joseph AM, Almutairi RE, Alrashidi WA, Almutairi RM, Theeban Y, Aldhuwayhi S et al. Factors Influencing Large Language Model Adoption among Dental Students: A Cross-Sectional Study. Scientific Reports. 2026 Apr. doi: 10.1038/s41598-026-47512-8. DOI: https://doi.org/10.1038/s41598-026-47512-8
Umer F, Batool I, Naved N. Innovation and Application of Large Language Models (LLMs) in Dentistry – A Scoping Review. British Dental Journal Open. 2024; 10(1): 90. doi: 10.1038/s41405-024-00277-6. DOI: https://doi.org/10.1038/s41405-024-00277-6
Najeeb M and Islam S. Artificial Intelligence (AI) in Restorative Dentistry: Current Trends and Future Prospects. BioMed Central Oral Health. 2025 Apr; 25(1): 592. doi: 10.1186/s12903-025-05989-1. DOI: https://doi.org/10.1186/s12903-025-05989-1
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Pakistan BioMedical Journal

This work is licensed under a Creative Commons Attribution 4.0 International License.
This is an open-access journal and all the published articles / items are distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. For comments editor@pakistanbmj.com

