Urgent.News

What's breaking now, across thousands of outlets.

Tech

How to Solve Hallucination (with RLCD)

LLMs, like large language models, provide confidence numbers based on their own "vibes" rather than precise measurements. For instance, when asked about weather forecasts, an LLM might confidently state a temperature and a 90% certainty level, even if the task of determining how that confidence number was reached proves impossible.

There are two main sources of uncertainty for an LLM: epistemic uncertainty (knowledge-based uncertainty) and aleatoric uncertainty (inherent randomness). The more historical data available, the lower the epistemic uncertainty; however, aleatoric uncertainty remains constant. This unpredictability is exemplified by asking a model to forecast the next day's temperature based on recent data, where a single day of history may not be sufficient to yield a reliable prediction.

Probability distributions can be employed to express uncertainty more accurately. Instead of providing a single temperature forecast, the model can generate a range of possible temperatures and assign a probability to each, resulting in a full range of probabilities that sum to 1. The distribution can then be compared to the actual distribution derived from historical data.

A calibrated model's forecast distribution aligns with the outcomes, such as having a 30% chance of temperatures within a specific range (e.g., 60-69°F) materializing in reality.

Reinforcement Learning from Correct Distribution (RLCD) is a method that teaches LLMs to generate calibrated probability forecasts. The process involves the model making a forecast and then creating several variations of that forecast with slightly different confidence levels. The next day's actual outcome is used to score each version, with higher scores awarded to versions that correctly predicted the outcome.

Over time, the model learns to favor the version that maximizes its score, thereby promoting honest self-assessment of its confidence levels. This training results in a model that not only makes accurate predictions but also communicates its level of certainty in those predictions.

One practical application of calibrated LLMs is in financial decision-making. For example, a model could analyze recent stock data, options flow, and news to generate a probability distribution over potential tomorrow's closing prices. The probabilities can then be used in conjunction with the Kelly criterion to determine an optimal investment strategy.

This approach allows investors to make informed decisions based on the model's calibrated confidence levels rather than relying on potentially misleading uncalibrated probabilities.

A real-world demonstration of RLCD's efficacy is presented in the case of NVDA stock. By providing the model with the last eight sessions of NVDA's actual OHLCV data, the model generates a calibrated probability distribution over potential closing prices. The distribution is then compared to the empirical distribution derived from the last 49 daily returns.

The calibrated distribution shows an overconfident shape where probability is concentrated in a single bucket, whereas the empirical distribution reveals a more spread-out outcome. When the calibrated probability is fed into the Kelly criterion, the resulting bet size is adjusted to reflect the true level of confidence in the forecast.

In this case, the calibrated distribution suggests a 0.62 chance of a price increase, corresponding to a Kelly fraction of 0.90, indicating that nine-tenths of the investment bankroll should be risked. However, since the distribution is not calibrated to the specific domain, the decision to invest nine-tenths of the bankroll may prove to be overly aggressive.

Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at robw.fyi →

More in Tech

I Gave My Proposal Agent Hindsight of Every Lost Bid

Aura Memory takes an RFP, works out what the client is asking for, writes an 11-section proposal, and then waits. When the deal closes, someone records the outcome: won, lost or pending, the factors…

  • Aura Memory analyzes RFPs to generate proposals.
  • Hindsight agent memory stores past outcomes for reference.
  • System aims to retain, recall, and reflect on proposals.

Detect a website's tech stack in bulk with Python (a Wappalyzer-style lookup)

I built this actor; it's a paid tool on Apify with a free trial credit. "What is this site built with?" is easy to answer for one site with a browser extension.

  • Website Tech Stack Detector tool detects tech stacks in bulk using Python
  • Tool performs HTTP requests, DNS lookups, TLS handshakes for analysis
  • Output JSON includes tech stack details with confidence scores and evidence

I Built a Customer Support Agent That Remembers 🤖

Customer support becomes frustrating when users have to explain the same problem again and again. So, I built a memory-enabled AI Customer Support Agent that can remember useful information from…

  • AI customer support agent remembers previous interactions
  • React frontend and FastAPI backend process user queries
  • Groq language model generates context-aware responses

MCP Transports That Still Matter: stdio vs Streamable HTTP (and Why SSE Is a Trap)

Attributed Chinese → English compile (not original authorship) Original title: MCP Server 开发入门:手把手写一个能跑的 Server,三种协议怎么选 Author: 晚安code URL: https://juejin.cn/post/7673880140422955046 Date: 2026-08-15…

  • MCP protocol separates model interface from byte movement
  • Avoid deprecated Streamable HTTP Server-Sent Events (SSE)
  • Stdio default for local hosts, Streamable HTTP for remote deployments

More from Monday 28 September →