I Sell Memory APIs. I'm Also Building the Benchmark. Here's How I'm Trying Not to Rig It.
Hey everyone. This time I'll go through what got me started on this benchmark, and the core of how it's actually built. It started from reading complaints, not from an idea. The same ones kept coming up: numbers a vendor publishes don't match numbers someone else measures, swapping the model that does the grading moves the results more than the gap between the systems being compared, and because…
The author of the memory API explains the process behind building a fair benchmark while being cautious not to manipulate the results. The benchmark originated from complaints about inconsistent numbers published by vendors. The benchmark focuses on fairness, with no affiliation to any vendor, and follows strict rules for submissions and verification.
The corpus contains 103,572 turns of conversation, with 1,991 sessions and approximately 1.9 million tokens. It runs at four different haystack sizes, and the key axes include 1,547 questions across 14 different categories. The benchmark also includes 500 false-memory probes, images with 50 photos, and questions in 10 different languages.
Latency measurements are conducted separately and subtracted from the overall latency to provide a fair assessment of system performance.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.