{
  "id": 12969756,
  "title": "Your Self-Hosted AI Assistant Keeps Forgetting You. Here Is How We Measured a Fix.",
  "url": "https://urgent.news/2026/10/08/your-self-hosted-ai-assistant-keeps-forgetting-you-here-is-how-we",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-08T22:28:26.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/neusoftware/your-self-hosted-ai-assistant-keeps-forgetting-you-here-is-how-we-measured-a-fix-18hi"
  },
  "original_language": "en",
  "account": "Your self-hosted AI assistant often forgets information over time. This is a common issue that many guides fail to address. The problem typically crops up after about a week of using the assistant. The assistant begins to ask again about details you provided in the early stages, forgets corrected project specifics, and even invents answers that were previously correct. This issue was experienced while building ScallopBot, an open-source self-hosted assistant. To understand how to fix it, the author measured the problem and the proposed solution against a published benchmark. There are three types of memory failures: storage failure, retrieval failure, and fabrication failure. Storage failure occurs when the fact was never written down anywhere. Retrieval failure happens when the assistant can't find the stored fact due to keyword or vector search mismatches. Fabrication failure is the worst type, where the model generates a plausible but incorrect answer when it can't find a relevant response. The fix depends on the specific failure. To address these issues, you can dump raw transcripts, grep them to locate the fact, and determine if it's a storage or retrieval problem. For a memory store that survives restarts, it's best to keep the entire memory in one file, such as using SQLite. Dates should be embedded within the memory, not as separate metadata. When dealing with contradictory facts, it's better to keep both and mark the older one as superseded rather than overwriting it. The retrieval process should involve two searches and a gate, using both BM25 keyword search and dense vector search to improve accuracy. The gate should return only results above a certain relevance threshold, and if none are found, the assistant should admit it doesn't know the answer. Regular maintenance, such as nightly consolidation cycles to merge near-duplicates and prune outdated information, is also recommended. In terms of performance, running this architecture against a benchmark showed an improvement in F1 score from 0.38 to 0.48. It also performed significantly better on adversarial questions with no correct answer stored (0.97 vs 0.77). However, these results should be considered as a direction rather than a definitive benchmark. The implementation is available under the MIT license, and the memory store is a single SQLite file that can be accessed and inspected.",
  "summary": "Your self-hosted AI assistant keeps forgetting you. Here is how we measured a fix. Every guide to running your own AI assistant covers the easy half: install Ollama, point a web UI at it, add an API key for a bigger model when you need one. None of them warn you about the hard half, which shows up around week three. The assistant starts forgetting things you told it in week one. Not in a dramatic…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}