Urgent.News

What's breaking now, across thousands of outlets.

Tech

Memory Support Agent: Building Customer Support That Remembers

Memory Support Agent: Building Customer Support That Remembers Introduction Customer support can become frustrating when customers have to explain the same problem every time they contact a support team. Most conversational AI systems can respond to a conversation, but persistent memory can make the experience much more useful. Instead of treating every interaction as completely new, an AI agent…

A support agent can provide accurate responses, yet still require customers to repeat themselves. To address this issue, the system incorporates long-term memory to maintain useful context across conversations while ensuring memory does not override account data that defines the true state. The application serves as a customer support console featuring chat, requests, orders, feedback, and account settings.

When a customer sends a message through the chat interface, the system retrieves relevant account context and generates an appropriate response. Attachments are stored alongside conversations, and every interaction is logged to provide historical context for subsequent exchanges. This process appears routine until the customer returns a week later.

Previous conversations might contain preferences, unsuccessful troubleshooting steps, or indications that an issue remains unresolved. Simply sending the entire transcript to a language model is an inefficient method to retrieve this information, as it increases prompt size, includes irrelevant turns, and complicates the distinction between customer statements and system knowledge about their account. To tackle this issue, the system treats context as distinct components with unique purposes:

* **Customer profile**: Holds stable preferences and identifying details.

* **Orders and requests**: Offers current account facts.

* **Recent chat history**: Preserves the immediate conversational thread.

* **Hindsight**: Offers durable, queryable memories from earlier interactions.

* **Historical support tickets**: Provides possible patterns, not specific facts about the customer.

The key engineering decision was to keep these sources separate until the moment a reply is generated. This approach enhances context richness without assuming equal authority for all retrieved text. The memory layer utilizes Hindsight on GitHub, with each customer having a unique bank. Retrieval occurs using the current message as the query, as demonstrated below:

```python

def _bank(customer_id):

return f"customer-{customer_id}"

def recall_memories(customer_id, query: str, limit: int = 6):

list[str]:

try:

result = _hindsight().recall(bank_id=_bank(customer_id), query=query)

items = getattr(result, "results", result) or []

return [str(getattr(item, "text", item)) for item in items][:limit]

except Exception:

return []

```

Two crucial factors influence this process. First, the customer identifier establishes memory boundaries. Second, retrieval is tailored to the current question, preventing the agent from randomly reiterating all stored information. The result limit ensures the long-term context remains bounded. The deployed application extracts this identifier from authenticated server-side state, rather than from any information supplied by the chat client.

Treating a memory system as a separate entity from the order service or request database is crucial for its effectiveness. Customer preferences, previous troubleshooting steps, or recurring explanations can significantly impact future exchanges. However, an order's current delivery status must come directly from the authoritative order record.

This separation is why the memory system is specifically designed to assist the agent, rather than treating a vector store as a transcript repository. The Hindsight documentation outlines the memory interface used by the application. The broader distinction between an agent's working context and its longer-lived memory is also covered in Vectorize's overview of agent memory.

Retrieval, not dumping, is the approach employed. The reply path assembles a concise view of the customer by retrieving memories alongside the new message, identifying relevant historical cases, and including only the most recent turns from the active conversation. Memories = recall_memories(customer[ customer_id ], user_msg) cases = similar_tickets(user_msg) messages = [{ role: "system", content: SYSTEM_PROMPT + "\n\n" + context }] messages += history[-8:] messages.append({ role: "user", content: user_msg }) return _llm(messages) This approach acknowledges that eight turns or six memories may not be universally suitable.

These bounds provide an inspectable context policy. Recent turns aid the model in understanding pronouns and handling follow-up questions, while Hindsight can contribute details from earlier sessions even if they fall outside the short retrieval window. The current message provides a concrete relevance signal for retrieval. The system also maintains separate searches for similar tickets.

The application utilizes TF-IDF to retrieve a few related support cases from a ticket dataset, offering general hints for suggested next steps. However, these cases should not be considered evidence about the customer in front of the system. A similar invoice dispute does not prove that the customer's invoice was duplicated. This distinction is enforced in system instructions, not through informal conventions.

"Use ONLY the customer facts, orders, requests, and memories given below. Never invent order numbers, refunds, dates, or policies." Personal memory can influence the agent's response, while account records establish what actually transpired. Historical cases can inform potential solutions but cannot fill gaps in a customer's record.

This separation is more beneficial than expecting a model to be cautious and hoping it correctly interprets every context source. A billing conversation example illustrates this approach: "I'm seeing the double charge again. Can you check?" The current message can retrieve earlier context related to billing. The customer profile provides the preferred contact method.

The request system can verify if a billing ticket is open. The latest turns maintain the follow-up continuity, ensuring a seamless customer experience.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at dev.to →

More in Tech

I deployed ONE Dockerized Notes App (Node.js + Postgres) locally + to Azure ACI! (CLOUD)

I Deployed Docker App on Azure — Part 1: Local + ACI (Working) Stack: Node.js 20 + Express + PostgreSQL 16 + Docker + Azure Container Instances Repo old: https://github.com/krisking7/lexszy-note New…

  • Dockerized Notes App deployed locally and to Azure Container Instances
  • App written in Node.js 20 with Express and PostgreSQL 16
  • Lessons learned about database readiness and ACI's ephemeral nature

Simulating Last Look: The Broker's Hold Window in 50 Lines of Python

Last look is a short window where the other side can accept or reject your order after seeing it. Here is a tiny simulation that shows how symmetric and asymmetric versions differ.

  • Simulates broker's last look hold window in Python
  • Compares symmetric and asymmetric last look scenarios
  • Reveals hidden asymmetries in execution process

How I Stopped Bioreactor Batch Losses Using Hindsight Memory

How I Stopped Bioreactor Batch Losses Using Hindsight Memory When a 500-liter bioreactor run experiences a sudden pH drop at 2 AM, standard LLM prompts offer textbook advice that wastes critical…

  • System uses LLM with hindsight memory to diagnose bioreactor issues
  • Retrieves past incident resolutions and anomaly signatures in seconds
  • Tested effectiveness with live pH drop and DO spike alert

Xcode 27: The Build Settings That Can Stop a Release

An unchanged Java app can stop building because a number inside its generated Xcode project is now too low. No new API call, no changed screen, just a deployment target the new SDK refuses to accept.

  • Xcode 27 updated build settings for compatibility
  • Removed three deprecated build hints in Xcode 27
  • Updated hints with new codename1.arg.ios settings

Apple Releases New AirPods Beta Firmware

Apple today released new beta firmware for the AirPods Pro 2, AirPods Pro 3 , AirPods 4, AirPods 5 , and AirPods Max 2 . The firmware has a build number of 9B5042a and is rolling out for developers…

More from Tuesday 29 September →