Validity and Reliability of Responses to Periodontology Questions by 4 Different Artificial Intelligence Chatbots as Public Information Sources
Cumhuriyet Dental Journal , vol.28, no.3, pp.390-396, 2025 (Scopus, TRDizin)
- Publication Type: Article / Article
- Volume: 28 Issue: 3
- Publication Date: 2025
- Doi Number: 10.7126/cumudj.1673333
- Journal Name: Cumhuriyet Dental Journal
- Journal Indexes: Scopus, Directory of Open Access Journals, TR DİZİN (ULAKBİM)
- Page Numbers: pp.390-396
- Ankara Yıldırım Beyazıt University Affiliated: Yes
Abstract
Objectives: To assess and check the validity and reliability of the answers given by ChatGpt-4o mini, Deepseek, Copilot and Gemini 1.5 flash daily chatbots to often seeked queries in the area of periodontology. Materials and Methods: Questions were selected from the most frequently asked patient questions by a periodontologist. Each question was asked to the chatbots three times. The answers (n = 240) were independently evaluated by two periodontologists on a Likert scale (5 = violently agree; 4 = agree; 3 = neutral; 2 = disagree; 1 = violently disagree). Disputes in scoring were removed through evidence-based negotiations. In evaluating the validity of the answers: Low threshold was determined as a score ≥4 for whole three answers; high threshold was determined as a score 5 for whole three answers. Fisher's exact test was performed to compare the validity of the answers among the chatbots. Cronbach's alpha was computed to evaluate the consistency and reliability of recurrent answers for each chatbot. Results: All four chatbots answered the questions. In the low-threshold validity test, ChatGpt had 100%, Deepseek and Copilot had 95%, Gemini had 65%. Gemini was significantly different from the others (p < 0.05). In the high-threshold validity test, ChatGpt had 80%, Deepseek had 75%, Copilot and Gemini were significantly lower at 5%. While there was no significant difference between ChatGpt and Deepseek (p > 0.05), both were significantly higher than Copilot and Gemini (p < 0.05). All four chatbots reached an acceptable level of reliability (Cronbach's alpha >0.7). Conclusions: ChatGpt and Deepseek provided more reliable information on periodontology-related topics than Copilot and Gemini.
Amaç: Periodontoloji alanında sık sorulan sorulara ChatGpt-4o mini, Deepseek, Copilot ve Gemini 1.5 flash günlük sohbet robotları tarafından verilen cevapların geçerliliğini ve güvenilirliğini değerlendirmek ve karşılaştırmaktır. Gereç ve Yöntemler: Sorular uzman bir periodontolog tarafından en çok sorulan hasta soruları arasından seçildi. Her soru, sohbet robotlarına üçer kez soruldu. Cevaplar (n = 240), iki periodontolog tarafından 5 puanlık Likert ölçeğiyle (5 = kesinlikle katılıyorum; 4 = katılıyorum; 3 = nötr; 2 = katılmıyorum; 1 = kesinlikle katılmıyorum) bağımsız olarak değerlendirildi. Puanlamadaki anlaşmazlıklar, kanıta dayalı tartışmalar yoluyla çözüldü. Yanıtların geçerliliği değerlendirilirken: Düşük eşik, üç cevabın tümü için puan ≥4 olarak belirlenirken; yüksek eşik, üç cevabın tümü için puan 5 olarak belirlendi. Sohbet robotları arasındaki cevapların geçerliliğini karşılaştırmak için Fisher'ın kesin testi yapıldı. Her sohbet robotu için tekrarlanan cevapların tutarlılığını ve güvenilirliği değerlendirmek için Cronbach'ın alfa değeri hesaplandı. Bulgular: Dört sohbet robotu da tüm sorulara cevap verdi. Düşük eşikli geçerlilik testinde ChatGpt %100, Deepseek ve Copilot %95, Gemini %65 geçerliliğe sahipti. Gemini diğerlerinden anlamlı farklıydı (p < 0.05). Yüksek eşikli geçerlilik testinde ChatGpt %80, Deepseek %75’di, Copilot ve Gemini %5 olarak önemli ölçüde düşüktü. ChatGpt ve Deepseek arasında anlamlı fark yokken (p > 0.05) her ikisi de Copilot ve Geminiden önemli ölçüde daha yüksekti (p < 0.05). Dört sohbet robotu da kabul edilebilir bir güvenilirlik düzeyine ulaştı (Cronbach'ın alfa değeri >0,7). Sonuçlar: ChatGpt ve Deepseek, Copilot ve Gemini’ye göre periodontoloji ile ilgili konularda daha güvenilir bilgiler sağladı.