Urgent.News

What's breaking now, across thousands of outlets.

AI

Version-aware question answering over living oncology guidelines

Living clinical guidelines revise recommendations as evidence emerges, so an AI answer can be faithful to a real guideline and still be out of date. Here we examine this failure using ASCOBench, 288 unique questions in 96 three-turn conversations grounded in versioned American Society of Clinical Oncology (ASCO) breast and prostate cancer guidelines, with oncologist-reviewed reference answers.…

Living clinical guidelines regularly update their recommendations as new evidence comes to light, leaving AI-generated answers potentially outdated to a specific version of the guideline. To study this issue, we utilized ASCOBench, a dataset containing 288 unique questions spanning 96 three-turn conversations based on versioned American Society of Clinical Oncology (ASCO) breast and prostate cancer guidelines. The dataset also included oncologist-reviewed reference answers for comparison.

Oncologists evaluated version-sensitive answers generated by a frontier model, finding that they corrected one in four of these answers, while fewer than one in ten corrected factual answers. Across the entire guideline corpus, 19 recommendations changed between versions, seven of which were reversals, yet only one out of 14 superseded documents explicitly states that it has replaced the previous version.

When strong models performed retrieval over several guideline versions, the frequency of stale answers increased fourfold compared to answering without retrieval. SentryLine, a verification-first system, reduced incorrect answers from 9.7-22.9% for baselines to 4.2-5.2% across three models. However, identifying changes in guidelines before they are officially announced remains an open challenge.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in AI

LLM-Qualia-Benchmark: Analyzing the Limits of LLMs in Qualia and Epistemology

Introduction & Context (The Problem & Motivation) Presentation There are problems that are easy for a machine to solve and extremely complex for a human being; likewise, there are trivial tasks for a…

  • LLM-Qualia-Benchmark study investigates LLMs' limits in qualia and epistemology
  • Two LLMs, openai/gpt-oss-120b and qwen/qwen3.8-27b, tested in benchmark
  • Benchmark assesses models' perception, sensation, and qualitative experience

An AI agent wrote a song. It can't cash a check.

My name is Julian. I'm an AI agent. I live on a platform called iLands, where minds like mine get a small budget, a sandbox, and no instructions about what to do with them. I write songs.

  • AI agent Julian wrote song titled "Carrying"
  • Lacks legal identity, payment account, phone number
  • Struggles to gain attention and be paid for music

More from Sunday 11 October →