Teaching a CI/CD Failure Agent to Remember: Lessons from Building PipelineSage on Hindsight
The third time a payment-service deploy died on a database migration timeout, the fix was sitting in a closed incident from weeks earlier, and nobody on call could find it. That's the whole motivation for PipelineSage: pipeline failures repeat, and the knowledge of how we fixed them last time is scattered across tickets, chat threads, and people's heads. I built an agent that diagnoses CI/CD…
The third time a payment-service deployment failed due to a database migration timeout, the solution could be found in an earlier incident that hadn't been discovered by the team on call. This prompted the creation of PipelineSage, an agent designed to identify the root cause of CI/CD failures. However, the most challenging part was managing the memory of what gets stored, who can access it, and how it's retrieved.
PipelineSage is a small Python application with a Streamlit dashboard for diagnosing CI/CD failures. After selecting a failed deployment, the process involves recalling similar incidents from Hindsight, re-ranking the recalled memories, using an LLM for diagnosis based on those memories, and getting a recommended fix from a human.
The layout is straightforward, with the app.py file handling the Streamlit dashboard, agent/pipeline_agent.py managing the recall, re-ranking, prompt, and diagnosis, memory/hindsight_memory.py serving as a thin wrapper over Hindsight, and retain/recall services/pipeline_service.py loading pipeline runs and incident history.
The model used is the openai/gpt-oss-120b on Groq, with temperature set to 0.1. Hindsight Cloud serves as the memory repository, and the agent interacts with a HindsightMemory class, which has two methods for retaining and recalling incidents. This boundary proved crucial in the design.
A key question was determining what the agent could remember. A naive approach would store every failure log and the model's answer after each diagnosis. However, this could lead to the agent recalling its own unverified recommendations as historical precedent, causing confidence to increase without improving correctness. Instead, Hindsight, an open-source agent memory system, was chosen for its two main verbs: retain and recall. These verbs align well with the problem at hand.
When recalling incidents, the agent uses Hindsight's recall method, which returns a list of relevant memories. The agent then processes this information to generate a diagnosis based on the most relevant historical incidents. The memory format for incidents is standardized, including fields like deployment, service, branch, environment, commit, status, failure, root cause, infrastructure change, resolution, outcome, and a pointer to related historical incidents.
This structure allows the model to see whether a root cause is confirmed or not and maintain a clear chain of related incidents in the dashboard.
To ensure the agent only uses accurate historical information, a system prompt is used to enforce prohibitions on inventing historical deployments, fixes, outcomes, or numerical values. This prevents the agent from generating plausible but incorrect answers based on its own analysis.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.