Urgent.News

What's breaking now, across thousands of outlets.

AI

LLM-powered automatic scoring tools for memory recall

Free recall of naturalistic stimuli, such as films and stories, reveals what people remember and how they organize it. Scoring such recall requires matching each recalled utterance to the encoded stimulus, and doing this manually often makes large studies impractical. Existing automated methods mostly return similarity or aggregate scores rather than explicit matches between recalled and stimulus…

An open-source, model-agnostic pipeline has been developed using large language models (LLMs) to automate scoring of memory recall tasks involving naturalistic stimuli like films and stories. Traditional manual scoring methods are impractical for large studies, as they become time-consuming and costly.

The new pipeline leverages LLMs to break down both the stimulus material and the participant's recall into smaller information units. It then matches these corresponding units between the two texts and generates a table of unit-to-unit matches with accuracy labels. This allows standard analysis scripts to compute standard recall measures.

The pipeline was tested with four different LLMs (Claude Sonnet 4.5, GPT 5.4, Gemini 2.5 Pro, and Gemini 2.5 Flash) on three different datasets. On typed recall of four short stories, the agreement between each model's matches and two trained raters was found to be close to the agreement between the raters themselves, with a phi coefficient ranging from .73 to .79.

A second run of the pipeline reproduced each model's matches with a phi coefficient ranging from .91 to .97, which was more closely aligned than the raters' agreement with each other. When analyzing spoken recall of 10 films, every model correctly attributed the recall to the correct film with high agreement.

However, agreement varied more across models when analyzing uncorrected speech-to-text recall of a television episode. The median cost per participant was under $1 for all models. These findings suggest that the pipeline can provide a reliable and cost-effective alternative to manual scoring of free recall tasks.

Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at biorxiv.org →

More in AI

More from Thursday 8 October →