Comparative performance and temporal variability of large language models on orthodontic questions from a national dental specialty examination
Scientific Reports, Published online: 23 August 2026; doi:10.1038/s41598-026-67407-y Comparative performance and temporal variability of large language models on orthodontic questions from a national dental specialty examination
A study conducted by researchers at Mersin University in Turkey examined the performance of three AI-based chatbots – ChatGPT-4o, ChatGPT-4.5, and Gemini 2.5 Pro – on orthodontic questions from the Turkish Dental Specialty Examination. The researchers administered 179 multiple-choice questions, spanning 18 examinations from 2012 to 2024, to the chatbots at two different time points in April and July 2025. The chatbots were tested under identical, standardized conditions.
At the first testing point (T1), ChatGPT-4.5 achieved the highest accuracy at 82.12%, while Gemini 2.5 Pro had the best accuracy at the second testing point (T2), with a score of 89.94%. The researchers found that Gemini 2.5 Pro demonstrated a significant improvement in accuracy between the two testing periods (p < 0.001), while ChatGPT-4.5 showed a smaller increase (p = 0.035).
Regardless of the model, all three systems performed significantly worse on visually oriented questions compared to text-based questions (p < 0.05). The study also identified that the accuracy of the large language models varied across different topic categories.
The varying performance of the chatbots across different time points and question types suggests the need for caution when interpreting repeated assessments conducted through evolving commercial interfaces. However, due to the limited number of visually oriented questions in this study, it is difficult to draw definitive conclusions about the performance of large language models in multimodal tasks.
Further research utilizing larger and more diverse visual datasets is needed to better understand their capabilities in this area.
Written by urgent.news from Scientific Reports's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.