Title: Giving On-Call Engineers a Memory: Building On Call Memory with Hindsight
Article Production incidents have a strange property: the same ones keep coming back. A database connection pool fills up in March, a different service leaks connections in August, and each time an engineer starts from zero. The knowledge exists somewhere, in a postmortem doc, a Slack thread, or the head of someone who left the company. It just isn't available at 3 a.m. when the pager goes off.…
Production incidents, such as database connection pool exhaustion and service connection leaks, repeatedly occur. Engineers often start from scratch without access to previous knowledge during critical on-call hours. To address this, our team developed OnCall Memory during the Hindsight hackathon. This incident response agent for Northwind Pay analyzes new alerts and recommends root causes and fixes.
The agent's key feature is its ability to remember past incidents, their resolutions, and contributing factors. Stateless assistants, like general-purpose LLMs, give generic advice without context, leading to repetitive responses on subsequent incidents. OnCall Memory is a FastAPI backend with a single-page frontend. When an engineer submits an alert, the backend runs analysis with and without memory, displaying both answers side by side.
Hindsight provides three operations: retain, recall, and reflect. During retain, the agent seeds a memory bank with 25 realistic incidents, each including dates, affected services, error logs, root causes, fix steps, resolution times, and fix effectiveness. The agent learns from engineer feedback during recall and identifies patterns through reflect.
This iterative process results in increasingly accurate recommendations as the agent accumulates experience.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.