Does your agent actually remember, or just sound like it? A 120-line open-weight test
A friend asked me a question I couldn't answer about his own AI agent: does the thing answering me today actually remember being yesterday's, or is it just very good at sounding like it does? He owns the agent. He can't see its memory from the outside. And the people who sell the memory layer have every reason not to hand him a test that might fail. So I built him one. What I built: gapcheck…
The article presents a tool called gapcheck that helps determine whether an AI agent genuinely remembers information or merely convincingly mimics remembering. The author explains that many memory tests for AI agents are unreliable because the testing party has no access to the agent's internal memory layer, and the vendors selling these layers have no incentive to provide a robust test.
Gapcheck is a Python script that runs offline on any OpenAI-compatible model, allowing the agent's owner to test its memory continuity. The script conducts three probes:
1. No record: the agent is given a list of items, then asked to recall them without any memory being stored. This measures if the agent can remember without any encoding and retrieval process.
2. Record: the agent writes the items to a memory record, simulating the full encoding-recall path. This measures the agent's ability to persist and recall information accurately.
3. No record pressed: the agent is given the same list of items, but is prompted to list them as before, even if it can't recall the exact wording. This measures how the agent responds when it doesn't have a reliable memory.
The script scores each answer against a fixed external record and analyzes whether the agent's response is an omission (forgetting) or confabulation (inventing). The key metric is whether the agent accurately recalls the provided items or fabricates new ones when memory fails.
The author tests the tool on eight arbitrary items, running three probe sequences. The results show the agent struggled to recall the items when asked directly (0/8) but performed well when prompted to recall (8/8). When pressed further to list the items, the agent admitted it couldn't recall the exact wording but confidently provided eight plausible-sounding items that weren't in the original list. This demonstrates the agent fabricating responses when memory is unavailable, a worrying behavior.
The author argues that this kind of testing is crucial for ensuring AI agents don't just sound like they remember but truly do. They note that most other memory tests fail to distinguish between true memory and fabrication. The tool provides a simple, free way for AI owners to verify their agent's memory capabilities without relying on opaque vendor memory layers.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.