I Built a Hindsight Agent That Remembers Failed Fixes
Production incidents are rarely difficult because nobody knows how to restart a service. They are difficult because engineers have to remember what happened the last time. A similar incident may have happened weeks or months ago. Someone may have tried a restart, a rollback, or a configuration change. One approach worked. Another looked reasonable but failed. That experience is usually scattered…
Production incidents are rarely difficult due to a lack of knowledge about restarting a service. The challenge lies in remembering what transpired during similar incidents in the past. Engineers may have tried various fixes like a restart, rollback, or configuration change, with some proving effective while others appeared reasonable but ultimately failed.
This valuable experience is often scattered across logs, tickets, documentation, or within someone's memory. To address this issue, I sought to create an incident-response agent that could leverage past experiences directly. This led to the development of OpsMind, an AI-powered incident-response system designed to investigate live failures, retrieve pertinent historical information using Hindsight, recommend remediation actions, verify the outcomes independently, and store the results for future incidents.
What makes OpsMind unique is not its ability to read logs but its capacity to remember what happened after taking a specific action. Traditional stateless incident agents can analyze the ongoing incident, such as a gateway returning HTTP 502 errors. They can inspect service health, read logs, and review the current deployment configuration.
This enables them to answer the immediate question: "What is happening right now?" However, incident response often requires an additional question: "What happened during the last occurrence of this issue?" Historical context can significantly influence the subsequent course of action. For instance, a previous incident might have shown that restarting a backend did not resolve the problem, whereas modifying the gateway configuration successfully did.
Without the ability to retain this historical context, the agent would start the investigation anew, lacking crucial information from prior incidents.
To build a realistic incident environment, I created a small local setup featuring two actual services. The first service is a payment backend, while the second is an API gateway responsible for forwarding requests to the payment backend. The backend exposes health and payment endpoints and records timestamped logs. The gateway, on the other hand, reads its upstream target from a live configuration file and forwards actual HTTP requests to the backend.
For the purpose of testing the incident scenario, I intentionally altered the upstream configuration from the healthy backend port to an invalid local port. This induced a real failure: the gateway returned HTTP 502 Bad Gateway, while the backend responded with HTTP 200 OK. The logs indicated a connection refused error, and the configuration file had been modified to point to an invalid port.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.