Temporally distinct reward and action prediction error signals during value learning and habit formation
Effective decision making in stochastic environments requires balancing flexible, value-based learning with a stabilising influence of habitual action selection. While dopamine-mediated reward prediction errors (RPEs) are a well-established component of value learning, the mechanisms underlying habit-like behaviour remain less clear. Here, we combined behavioural analysis, computational…
Effective decision making in uncertain situations necessitates a balance between adaptable value-based learning and a stabilizing effect of habitual action selection. Dopamine-mediated reward prediction errors (RPEs) are known to play a role in value learning, but the mechanisms behind habit-like behavior are less understood. In this study, researchers examined the decision-making process of mice engaged in a probabilistic choice task, in which the actions taken were not immediately tied to the outcome on each trial.
The team found that choice behavior was best explained by a model that integrated value-based, habitual, and risk-sensitive components, each of which was updated by separate reward- and action-related learning signals.
The researchers observed that dopamine activity in the dorsolateral striatum carried not only RPE-like signals when making a choice and receiving a result, but also temporally distinct action prediction errors (APEs) following the completion of a choice. These findings support a framework in which dopamine signals in the dorsolateral striatum carry parallel but distinct reward- and action-related learning signals to facilitate value- and habit-based processes.
Written by urgent.news from bioRxiv's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.