{
  "id": 7877809,
  "title": "Impaired Reinforcement Learning Underlying Explore-Exploit Decision Making in Theft Recidivists",
  "url": "https://urgent.news/2026/09/16/impaired-reinforcement-learning-underlying-explore-exploit-decision",
  "topic": "science",
  "section": "Science",
  "published": "2026-09-16T00:00:00.000Z",
  "source": {
    "name": "bioRxiv",
    "slug": "biorxiv",
    "url": "https://www.biorxiv.org/content/10.64898/2026.09.11.750836v1?rss=1"
  },
  "original_language": "en",
  "account": "A recent study explored the behavioral and neurobiological underpinnings of recurrent theft among theft recidivists. Researchers utilized a 4-arm bandit task, while simultaneously monitoring prefrontal cortex hemodynamics via functional near-infrared spectroscopy (fNIRS). The findings suggest a significant impairment in reinforcement learning mechanisms in non-kleptomanic theft recidivists compared to control participants and kleptomanic offenders.\n\nNon-kleptomanic theft recidivists (TR-K) accumulated substantially higher cumulative regret and made fewer optimal choices than both control individuals without criminal records (CT) and kleptomanic offenders (TR+K). The model-based analysis revealed that Q-learning with decay model best fit the observed data. The parameter extraction demonstrated a lower learning rate in the TR-K group compared to the CT and TR+K groups, indicating a deficit in updating action values following environmental feedback.\n\nThe fNIRS tracking of trial-by-trial latent reinforcement variables showed that PFC activity was modulated by these variables. However, the group differences were characterized by static baseline hemodynamic shifts rather than rewirings of value-tracking neural circuits. In conclusion, these results suggest that impaired reinforcement learning mechanisms in non-kleptomanic theft recidivists challenge the effectiveness of current punitive deterrence models.",
  "summary": "Larceny imposes profound societal and economic burdens; however, punitive judicial measures frequently fail to deter recidivism. The neurobehavioral mechanisms driving habitual offending, whether instrumental or kleptomanic, in theft recidivists remain poorly understood. In this study, we investigated explore-exploit decision-making and underlying reinforcement learning architectures in theft…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}