{
  "id": 10725353,
  "title": "I Built a Support Agent That Remembers Customers Between Chats",
  "url": "https://urgent.news/2026/09/29/i-built-a-support-agent-that-remembers-customers-between-chats",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-29T16:05:15.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/manasa_penchala/i-built-a-support-agent-that-remembers-customers-between-chats-3o16"
  },
  "original_language": "en",
  "account": "A support agent can provide accurate responses, yet still require customers to repeat themselves. To address this issue, the system incorporates long-term memory to maintain useful context across conversations while ensuring memory does not override account data that defines the true state. The application serves as a customer support console featuring chat, requests, orders, feedback, and account settings. When a customer sends a message through the chat interface, the system retrieves relevant account context and generates an appropriate response. Attachments are stored alongside conversations, and every interaction is logged to provide historical context for subsequent exchanges. This process appears routine until the customer returns a week later. Previous conversations might contain preferences, unsuccessful troubleshooting steps, or indications that an issue remains unresolved. Simply sending the entire transcript to a language model is an inefficient method to retrieve this information, as it increases prompt size, includes irrelevant turns, and complicates the distinction between customer statements and system knowledge about their account. To tackle this issue, the system treats context as distinct components with unique purposes:\n\n* **Customer profile**: Holds stable preferences and identifying details.\n* **Orders and requests**: Offers current account facts.\n* **Recent chat history**: Preserves the immediate conversational thread.\n* **Hindsight**: Offers durable, queryable memories from earlier interactions.\n* **Historical support tickets**: Provides possible patterns, not specific facts about the customer.\n\nThe key engineering decision was to keep these sources separate until the moment a reply is generated. This approach enhances context richness without assuming equal authority for all retrieved text. The memory layer utilizes Hindsight on GitHub, with each customer having a unique bank. Retrieval occurs using the current message as the query, as demonstrated below:\n\n```python\ndef _bank(customer_id):\nreturn f\"customer-{customer_id}\"\n\ndef recall_memories(customer_id, query: str, limit: int = 6):\nlist[str]:\ntry:\nresult = _hindsight().recall(bank_id=_bank(customer_id), query=query)\nitems = getattr(result, \"results\", result) or []\nreturn [str(getattr(item, \"text\", item)) for item in items][:limit]\nexcept Exception:\nreturn []\n```\n\nTwo crucial factors influence this process. First, the customer identifier establishes memory boundaries. Second, retrieval is tailored to the current question, preventing the agent from randomly reiterating all stored information. The result limit ensures the long-term context remains bounded. The deployed application extracts this identifier from authenticated server-side state, rather than from any information supplied by the chat client. Treating a memory system as a separate entity from the order service or request database is crucial for its effectiveness. Customer preferences, previous troubleshooting steps, or recurring explanations can significantly impact future exchanges. However, an order's current delivery status must come directly from the authoritative order record. This separation is why the memory system is specifically designed to assist the agent, rather than treating a vector store as a transcript repository. The Hindsight documentation outlines the memory interface used by the application. The broader distinction between an agent's working context and its longer-lived memory is also covered in Vectorize's overview of agent memory. Retrieval, not dumping, is the approach employed. The reply path assembles a concise view of the customer by retrieving memories alongside the new message, identifying relevant historical cases, and including only the most recent turns from the active conversation. Memories = recall_memories(customer[ customer_id ], user_msg) cases = similar_tickets(user_msg) messages = [{ role: \"system\", content: SYSTEM_PROMPT + \"\\n\\n\" + context }] messages += history[-8:] messages.append({ role: \"user\", content: user_msg }) return _llm(messages) This approach acknowledges that eight turns or six memories may not be universally suitable. These bounds provide an inspectable context policy. Recent turns aid the model in understanding pronouns and handling follow-up questions, while Hindsight can contribute details from earlier sessions even if they fall outside the short retrieval window. The current message provides a concrete relevance signal for retrieval. The system also maintains separate searches for similar tickets. The application utilizes TF-IDF to retrieve a few related support cases from a ticket dataset, offering general hints for suggested next steps. However, these cases should not be considered evidence about the customer in front of the system. A similar invoice dispute does not prove that the customer's invoice was duplicated. This distinction is enforced in system instructions, not through informal conventions. \"Use ONLY the customer facts, orders, requests, and memories given below. Never invent order numbers, refunds, dates, or policies.\" Personal memory can influence the agent's response, while account records establish what actually transpired. Historical cases can inform potential solutions but cannot fill gaps in a customer's record. This separation is more beneficial than expecting a model to be cautious and hoping it correctly interprets every context source. A billing conversation example illustrates this approach: \"I'm seeing the double charge again. Can you check?\" The current message can retrieve earlier context related to billing. The customer profile provides the preferred contact method. The request system can verify if a billing ticket is open. The latest turns maintain the follow-up continuity, ensuring a seamless customer experience.",
  "summary": "A support agent can give the right answer and still make a customer repeat themselves. I built this system around a simple idea: use long-term memory to carry useful context between conversations, but never let memory outrank the account data that determines what is actually true. The Context Problem The application is a customer support console with chat, requests, orders, feedback, and account…",
  "key_points": [
    "Memory layer uses Hindsight with unique bank per customer for retrieval.",
    "Retrieval is tailored to current query, preventing random reiteration of stored information."
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "Dev.to",
        "title": "I Built a Customer Support Agent That Remembers With Hindsight",
        "url": "https://urgent.news/2026/09/29/i-built-a-customer-support-agent-that-remembers-with-hindsight",
        "published": "2026-09-29T17:28:08.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}