{
  "id": 12915909,
  "title": "LLM-powered automatic scoring tools for memory recall",
  "url": "https://urgent.news/2026/10/08/llm-powered-automatic-scoring-tools-for-memory-recall",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-08T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.10.01.755994v1?rss=1"
  },
  "original_language": "en",
  "account": "An open-source, model-agnostic pipeline has been developed using large language models (LLMs) to automate scoring of memory recall tasks involving naturalistic stimuli like films and stories. Traditional manual scoring methods are impractical for large studies, as they become time-consuming and costly.\n\nThe new pipeline leverages LLMs to break down both the stimulus material and the participant's recall into smaller information units. It then matches these corresponding units between the two texts and generates a table of unit-to-unit matches with accuracy labels. This allows standard analysis scripts to compute standard recall measures.\n\nThe pipeline was tested with four different LLMs (Claude Sonnet 4.5, GPT 5.4, Gemini 2.5 Pro, and Gemini 2.5 Flash) on three different datasets. On typed recall of four short stories, the agreement between each model's matches and two trained raters was found to be close to the agreement between the raters themselves, with a phi coefficient ranging from .73 to .79.\n\nA second run of the pipeline reproduced each model's matches with a phi coefficient ranging from .91 to .97, which was more closely aligned than the raters' agreement with each other. When analyzing spoken recall of 10 films, every model correctly attributed the recall to the correct film with high agreement.\n\nHowever, agreement varied more across models when analyzing uncorrected speech-to-text recall of a television episode. The median cost per participant was under $1 for all models. These findings suggest that the pipeline can provide a reliable and cost-effective alternative to manual scoring of free recall tasks.",
  "summary": "Free recall of naturalistic stimuli, such as films and stories, reveals what people remember and how they organize it. Scoring such recall requires matching each recalled utterance to the encoded stimulus, and doing this manually often makes large studies impractical. Existing automated methods mostly return similarity or aggregate scores rather than explicit matches between recalled and stimulus…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}