Urgent.News

What's breaking now, across thousands of outlets.

AI

How I fixed LLM counting hallucinations using Hindsight facts

The first time my support agent told me a customer had contacted us "twice" when the record showed four separate contacts, I assumed I had a retrieval bug. I didn't. The memory was fine. I had asked a language model to count, and it did what language models do when you ask them to count: it produced a confident, plausible number. This is the story of how I stopped asking the model to count, and…

A customer support memory agent was developed to streamline customer interactions across chat, email, and phone channels. When a support agent reported discrepancies in the number of customer contacts, it became apparent that a language model was being asked to count interactions, leading to inaccurate results. The issue stemmed from the model's difficulty in determining whether multiple messages referred to the same problem or not.

To address this problem, the solution involved storing the interactions as structured facts and performing the counting in Python. This approach eliminated the need for the model to make judgement calls and provided an intermediate value that could be logged, tested, and verified. By retaining the facts in Hindsight, the agent could serve different purposes for the summarise and escalate endpoints, using the same history without the need for separate databases.

This change eliminated the hallucinations caused by the model's counting and provided a more reliable method for determining if a customer had contacted the support team multiple times about the same unresolved issue.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Tuesday 29 September →