{
  "id": 10474345,
  "title": "How to Solve Hallucination (with RLCD)",
  "url": "https://urgent.news/2026/09/28/how-to-solve-hallucination-with-rlcd",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-28T15:04:58.000Z",
  "source": {
    "name": "Lobsters",
    "slug": "lobsters",
    "url": "https://www.robw.fyi/2026/09/28/how-to-solve-hallucination/"
  },
  "original_language": "en",
  "account": "LLMs, like large language models, provide confidence numbers based on their own \"vibes\" rather than precise measurements. For instance, when asked about weather forecasts, an LLM might confidently state a temperature and a 90% certainty level, even if the task of determining how that confidence number was reached proves impossible.\n\nThere are two main sources of uncertainty for an LLM: epistemic uncertainty (knowledge-based uncertainty) and aleatoric uncertainty (inherent randomness). The more historical data available, the lower the epistemic uncertainty; however, aleatoric uncertainty remains constant. This unpredictability is exemplified by asking a model to forecast the next day's temperature based on recent data, where a single day of history may not be sufficient to yield a reliable prediction.\n\nProbability distributions can be employed to express uncertainty more accurately. Instead of providing a single temperature forecast, the model can generate a range of possible temperatures and assign a probability to each, resulting in a full range of probabilities that sum to 1. The distribution can then be compared to the actual distribution derived from historical data. A calibrated model's forecast distribution aligns with the outcomes, such as having a 30% chance of temperatures within a specific range (e.g., 60-69°F) materializing in reality.\n\nReinforcement Learning from Correct Distribution (RLCD) is a method that teaches LLMs to generate calibrated probability forecasts. The process involves the model making a forecast and then creating several variations of that forecast with slightly different confidence levels. The next day's actual outcome is used to score each version, with higher scores awarded to versions that correctly predicted the outcome. Over time, the model learns to favor the version that maximizes its score, thereby promoting honest self-assessment of its confidence levels. This training results in a model that not only makes accurate predictions but also communicates its level of certainty in those predictions.\n\nOne practical application of calibrated LLMs is in financial decision-making. For example, a model could analyze recent stock data, options flow, and news to generate a probability distribution over potential tomorrow's closing prices. The probabilities can then be used in conjunction with the Kelly criterion to determine an optimal investment strategy. This approach allows investors to make informed decisions based on the model's calibrated confidence levels rather than relying on potentially misleading uncalibrated probabilities.\n\nA real-world demonstration of RLCD's efficacy is presented in the case of NVDA stock. By providing the model with the last eight sessions of NVDA's actual OHLCV data, the model generates a calibrated probability distribution over potential closing prices. The distribution is then compared to the empirical distribution derived from the last 49 daily returns. The calibrated distribution shows an overconfident shape where probability is concentrated in a single bucket, whereas the empirical distribution reveals a more spread-out outcome. When the calibrated probability is fed into the Kelly criterion, the resulting bet size is adjusted to reflect the true level of confidence in the forecast. In this case, the calibrated distribution suggests a 0.62 chance of a price increase, corresponding to a Kelly fraction of 0.90, indicating that nine-tenths of the investment bankroll should be risked. However, since the distribution is not calibrated to the specific domain, the decision to invest nine-tenths of the bankroll may prove to be overly aggressive.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}