Abstract

Objective: This study aimed to explore common patient inquiries about Cone-Beam Computed Tomography (CBCT) and systematically assess the accuracy, quality, usefulness, and readability of outputs from four AI-based conversational platforms (ChatGPT-3.5, DeepSeek-V3.1, Gemini 1.5 Flash, and Microsoft Copilot Free).

Materials and Methods: Common questions about CBCT were gathered from online sources and expert contributions and submitted to ChatGPT-3.5, DeepSeek, Gemini, and Copilot under standardized conditions. Outputs were independently evaluated by three specialists using CLEAR, mGQS, accuracy, usefulness, DISCERN, and readability metrics (FRE and FKGL).

Results: Significant differences were observed among AI platforms regarding CLEAR scores (p < 0.05), with Gemini showing higher values than ChatGPT (p = 0.016). Across all four platforms, strong to a very strong positive correlations were found among CLEAR, mGQS, and Accuracy scores (r ≥ 0.77, p < 0.001), while these measures were strongly to very strongly and inversely correlated with Usefulness (r ≤ −0.77, p < 0.001). Flesch Reading Ease and Flesch-Kincaid Grade Level demonstrated a very strong negative correlations across platforms (r ranging from −0.85 to −0.93, p < 0.05). Within the Gemini group, Flesch-Kincaid Grade Level showed a moderate positive association with DISCERN scores (r = 0.496, p = 0.043).

Conclusions: Gemini and DeepSeek produced more accurate and clearer responses. The readability of all AI systems was below the recommended level for patient education. These systems may support patient education but require expert validation.

Keywords: CBCT, artificial intelligence, chatbots, patient education, readability

Copyright and license

How to cite

1.
Özarslantürk S, Ceylan Şen S, Saraç Atagün Ö, Gaş S, Paksoy T, Çardakçı Bahar Ş. Evaluation of artificial intelligence chatbots in informing patients about CBCT: A comparative study. Northwestern Med J. 2026;6(3):192-20. https://doi.org/10.54307/NWMJ.2026.229