Urgent.News

What's breaking now, across thousands of outlets.

AI

Comparative performance and temporal variability of large language models on orthodontic questions from a national dental specialty examination

Scientific Reports, Published online: 23 August 2026; doi:10.1038/s41598-026-67407-y Comparative performance and temporal variability of large language models on orthodontic questions from a national dental specialty examination

A study conducted by researchers at Mersin University in Turkey examined the performance of three AI-based chatbots – ChatGPT-4o, ChatGPT-4.5, and Gemini 2.5 Pro – on orthodontic questions from the Turkish Dental Specialty Examination. The researchers administered 179 multiple-choice questions, spanning 18 examinations from 2012 to 2024, to the chatbots at two different time points in April and July 2025. The chatbots were tested under identical, standardized conditions.

At the first testing point (T1), ChatGPT-4.5 achieved the highest accuracy at 82.12%, while Gemini 2.5 Pro had the best accuracy at the second testing point (T2), with a score of 89.94%. The researchers found that Gemini 2.5 Pro demonstrated a significant improvement in accuracy between the two testing periods (p < 0.001), while ChatGPT-4.5 showed a smaller increase (p = 0.035).

Regardless of the model, all three systems performed significantly worse on visually oriented questions compared to text-based questions (p < 0.05). The study also identified that the accuracy of the large language models varied across different topic categories.

The varying performance of the chatbots across different time points and question types suggests the need for caution when interpreting repeated assessments conducted through evolving commercial interfaces. However, due to the limited number of visually oriented questions in this study, it is difficult to draw definitive conclusions about the performance of large language models in multimodal tasks.

Further research utilizing larger and more diverse visual datasets is needed to better understand their capabilities in this area.

Written by urgent.news from Scientific Reports's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at nature.com →

More in AI

Superpowers for Coding Agents: Turn Vague Requests Into Tested Changes

The fastest way for an AI coding agent to create expensive work is to start coding too soon. A request arrives, the agent infers the missing requirements, selects an architecture, edits several files…

  • Superpowers methodology introduces gates to prevent costly coding mistakes.
  • Agent recognizes need to create something, then understands user intent.
  • Testing and subagent-driven development ensure final product aligns with user intent.

Bette Midler Flames Chart-Topping AI Song as ‘Cultural Genocide’

The legendary entertainer is airing her grievances over AI-generated artist IngaRose's viral hit "Celebrate Me" The post Bette Midler Flames Chart-Topping AI Song as ‘Cultural Genocide’ appeared first…

  • Bette Midler calls AI-generated song 'cultural genocide'
  • Singer criticizes lack of compensation for artists
  • Tyrese Gibson defends AI, Midler's stance resonates with others

Lets talk about llms

Differently. (ChatGPT and LLMs—a friendly reminder) (Assume everything is possible—another friendly reminder.) LLMs are great. LLMs are easy and efficient. Yeah? But for how long?

  • LLMs generate human-like text efficiently but pose risks.
  • Scraped content could harm individuals or properties.
  • Over-reliance on LLMs may lead to misinformation.

More from Sunday 23 August →