Urgent.News

What's breaking now, across thousands of outlets.

Tech

A new dual-process theory solves the mystery of dopamine ramps

A new computational model explains why dopamine levels gradually rise as animals approach a reward. By combining fast mental maps with slow habit formation, the dual-process theory resolves a long-standing puzzle about how the brain updates its expectations.

A new dual-process theory solves the mystery of dopamine ramps

In a groundbreaking study, researchers from the University of Oxford have unveiled a novel computational model that elucidates the enigmatic phenomenon of dopamine ramps. By integrating two distinct learning processes, the model demonstrates how the brain adeptly updates its expectations, thereby resolving a longstanding enigma in neuroscience. The findings have been published in the esteemed journal eLife.

Dopamine, a neurotransmitter pivotal to learning, motivation, and movement, has traditionally been perceived as a marker for reward prediction error—i.e., the discrepancy between an actual outcome and its anticipated value. When an unexpected reward is encountered, dopamine neurons fire to signal a positive error, enabling the brain to recalibrate its stored expectations in a region called the striatum.

However, experimental measurements during spatial navigation tasks have consistently revealed a counterintuitive pattern: as animals approach a known, predictable reward, their dopamine levels exhibit a progressive increase in a smooth gradient.

This observation posed a significant challenge to conventional theoretical frameworks, which struggle to account for why a prediction error would escalate as the goal nears. In response to this enigma, researchers Luke Priestley and Thomas Akam embarked on the development of a computational model that harmonizes two separate learning mechanisms operating concurrently.

The first process is a conventional, slow-learning system that relies on cached values lodged in the basal ganglia. The second process is a rapid, adaptable system that autonomously infers values using an internal map or world model, which is believed to reside within the brain's frontal cortex.

According to Priestley and Akam, these two systems engage in a highly specific interaction to generate dopamine ramps. When calculating a reward prediction error, the brain juxtaposes its current prediction against a novel update target. In their model, the fast, inferred values merely shape the update target, while the current prediction is solely derived from the slow, cached values.

As the fast system is already cognizant of the reward's proximity, while the slow system is still assimilating the information, the gap between the update target and the prediction widens as the goal draws nearer. This escalating gap engenders the gradual rise in dopamine.

To validate their model, Priestley and Akam subjected it to rigorous testing within a simulated linear track environment. They juxtaposed it against a standard model and an alternative version where inferred values influenced both the prediction and the update target. The asymmetrical model outperformed its counterparts by learning the true value of the environment more efficiently and successfully generating the ramping dopamine signals that eluded the standard models.

Subsequently, the researchers simulated an environment in which an artificial agent traversed between high and low rewards across thousands of trials. They modeled a previously documented experiment that demonstrated dopamine ramps in mice diminishing gradually after extensive training. The simulated agent mirrored this long-term decline. Moreover, as the slow-learning cached values converged with the fast-learning inferred values, the gap between them contracted, resulting in the flattening of the ramps over time.

The model also accounted for dopamine behavior in novel environments. In biological experiments, animals do not display dopamine ramps during their inaugural exploration of a new maze; however, the ramps emerge swiftly after a few successes. The simulated agents exhibited this very rapid onset, corroborating how the fast-learning internal map rapidly modulates the prediction error.

The researchers then applied their model to a grid-like environment with multiple pathways leading to a single destination. In real-world experiments, altering the quantity of reward at a specific location instantaneously modifies the dopamine ramp on the subsequent attempt, irrespective of the animal's route. The dual-process model successfully reproduced this global updating behavior, as the fast-learning system employs a flexible mental map to immediately apply the new reward information to all feasible paths leading to the goal.

To elucidate how unexpected events affect dopamine, the team simulated virtual reality experiments in which animals were suddenly transported closer to a goal or compelled to move at varying speeds. In the simulation, teleports induced abrupt spikes in the simulated dopamine signal, with the spike's magnitude contingent on the proximity of the agent to the reward.

Altering the agent's speed modified the steepness of the ramp, findings that aligned with actual biological recordings, thereby substantiating the notion that dopamine tracks transient changes in expected value.

Lastly, the researchers modeled spatial uncertainty by simulating a virtual reality task wherein the environment progressively darkened. In real-world animal experiments, this darkening incites dopamine levels to rise in a hump-shaped pattern rather than a steady ramp. The simulated agent produced these very same shapes, with the agent's reduced certainty about its exact location causing the prediction error to wane before attaining the goal as the visual environment darkened.

Written by urgent.news from PsyPost's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at psypost.org →

More in Tech

Eran’s Achilles heel; the selectors

by Rex Clementine The Cricket Transformation Committee headed by Eran Wickramaratne has received much praise for the manner in which it has gone about its business since being put in charge earlier…

More from Wednesday 12 August →