{
  "id": 281015,
  "title": "Actor-Critic Methods: Two Networks, One Loop",
  "url": "https://urgent.news/2026/08/07/actor-critic-methods-two-networks-one-loop",
  "topic": "culture",
  "section": "Culture",
  "published": "2026-08-07T21:18:57.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/multigrid/actor-critic-methods-two-networks-one-loop-4n75"
  },
  "original_language": "en",
  "account": "Actor-critic methods involve simultaneously training two components: a policy that determines actions, and a value function that predicts the quality of those actions. The value function serves as a guide for the policy, indicating which outcomes were better than expected. The primary challenge for the critic lies in estimating the return, a value that heavily depends on the length of the episode and the randomness of the environment.\n\nThe actor-critic approach mitigates the issue of high variance in sample requirements by employing a critic to address the shortcomings of the raw return value. The critic aims to predict the expected discounted return from a state under the current policy. However, the critic must be careful not to confuse a reward predictor with its actual function.\n\nThe core of the critic's task is to estimate V(s), which represents the expected discounted return from state s under the current policy. It is crucial to note that V(s) is not a reward predictor but rather a prediction of the sum of all future discounted rewards. Additionally, the critic's predictions are tied to the policy and not the environment itself. This leads to the interaction between learning rates, as improving the actor requires adjusting the critic's values accordingly.\n\nThe actor-critic architecture is often implemented using a single network with two heads sharing a trunk, which reduces computational requirements but introduces a trade-off. Unlike Q-learning, where the value function guides the policy through an arg-max operation, the actor-critic setup involves the actor making decisions and the critic providing feedback.\n\nTo illustrate the advantages of this method, consider a three-step episode with rewards only at the end, a discount factor (gamma) of 0.99, and a critic that initially predicts values for each state. By calculating the one-step deltas, we observe that the advantages are lower in variance but biased, while the full Monte Carlo returns are unbiased but high variance. The advantage of an action is the difference between the outcome and the critic's expectations, and it can significantly impact the learning process.\n\nThe balance between the two extremes is achieved through the use of a single parameter in the generalized advantage estimation formula. By adjusting this parameter, researchers can control the extent to which the critic's values influence the learning process. The default value of 0.95 has been found to work well, as it effectively discounts the credit chain by about five percent per step and provides an effective credit horizon of roughly 17 steps.\n\nThe actor-critic loop involves running the current policy for a batch of steps, recording states, actions, rewards, and the critic's value at each step. The advantages are computed backwards through the batch, and value targets are calculated as the sum of the advantages and the critic's value for each state. The actor is then updated using the policy gradient, scaling the advantages as the scaling term. Meanwhile, the critic is updated by regressing the value targets using mean squared error.\n\nAn essential aspect of the actor-critic method is the application of advantage normalization, where the batch mean and standard deviation are subtracted and divided, respectively. This normalization ensures that the effective learning rate remains independent of the reward scale, allowing researchers to change the reward scale without significantly affecting the training dynamics.\n\nHowever, the actor-critic method is not without its challenges. One common issue is when the critic lags behind the actor, leading to stale expectations and potentially causing the policy to improve in the wrong direction. The symptom of this failure is an increase in return, followed by oscillations or collapses. To address this issue, it is often beneficial to update the critic more frequently than the actor or to reduce the actor's learning rate.",
  "summary": "An actor-critic method trains two things at once: a policy that acts, and a value function that predicts how well things are going. The value function is not there to choose actions. It is there to tell the policy which of its results were better than expected. The problem the critic is hired to solve A policy gradient scales each action’s gradient by the return that followed it. That return is a…",
  "key_points": [
    "Actor-critic methods train two networks: policy and value function.",
    "Critic estimates expected discounted return, not reward predictor.",
    "Advantage normalization stabilizes learning rate across reward scales."
  ],
  "editors_take": null,
  "illustration": "https://urgent.news/ill/281015.png",
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}