From Email Export to Agent Memory: The Pipeline I Actually Wanted
From Email Export to Agent Memory: The Pipeline I Actually Wanted The interesting part of an AI application is often not the model call. It is the pipeline that decides what the model gets to remember. I built Waada as a deal-continuity system, and the engineering problem that kept surfacing was straightforward: how do I turn messy account history into durable, queryable memory without making…
The architecture begins with accepting exported sales history from multiple sources, including .eml emails, Slack day JSON text, Markdown, VTT transcripts, audio, and CRM JSON. The application, built with TypeScript and pnpm workspace, is composed of various packages, each handling different aspects of the pipeline. The core model is a Zod-backed Interaction, which normalizes all incoming data regardless of its original source format.
After normalization, the data is ingested and validated by the ingest() function, which ensures idempotency and sorts interactions by date. It also checks for existing data in the .waada/manifest.json file to avoid duplicating information. The Hindsight document ID provides a stable identity for each interaction, which is crucial because sales exports are often re-exported, and users may upload the same data multiple times.
The retained interactions are stored in Hindsight, a durable memory layer isolated in packages/core/src/memory/. Each account maps to a bank derived from its slug, and the Hindsight adapter retains content, interaction date, context, document ID, and metadata for each interaction. Hindsight's job is not to decide what a commitment means but to provide the durable account history for the agent to retrieve evidence.
Once the data is retained, the agent performs different retrieval jobs based on the operation. Commitment tasks search for promise-related evidence, landmine extractors look for objections and resolved agreements, the brief recalls stakeholder and recent-change evidence, and the Ask function uses the user's exact question. This purpose-specific query is performed through Hindsight recall, which returns a list of MemoryHits containing the relevant evidence.
The commitment ledger is a retrieval pipeline that performs multiple promise-oriented queries, deduplicates evidence, chunks it, runs structured extraction, and merges duplicates. The resulting Commitment contains the deliverable, parties, date, optional due date, status, evidence, and source. The status is derived based on the open/closed and unclear/delivered state of each item, ensuring that stale open values are not permanently stored.
Finally, the agent generates the response using a bounded evidence budget of 5,000 input tokens, approximated conservatively through character counts. Waada uses evidence caps and a shared prompt budget to ensure the LLM receives only the necessary information for its response.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.