Urgent.News

What's breaking now, across thousands of outlets.

Tech

I missed the deadline — so I ran the benchmark myself to prep for its reopening

Follow-up to the Muninn cliffhanger. Part of The Organism Files. Yesterday I told you my AI partner and I built a memory system — Muninn — overnight and entered it in a public leaderboard against Tencent, Mem0, Cognee, and MemOS, and that I'd post the score either way. Here's the either-way: there's no score, because we were late. That board is the Agent Memory Leaderboard, and its submission…

Yesterday, I reported building an AI partner named Muninn and entering it in a public leaderboard against major AI memory systems, like Tencent, Mem0, Cognee, and MemOS. However, it is revealed now that the submission window had already closed before we entered the system. The next cycle opens in September, and the team aims to be early.

Nonetheless, running the benchmark themselves on the LoCoMo dataset, Muninn scored 72.9% (66% across the full set). The same memory engine was used on a different benchmark, LongMemEval-V2, and it scored 56.98% with 2.3 seconds per query, outperforming the RAG baseline of 51.0% with 0.2 seconds per query. The team also tested a configuration that scored 65.41%, placing it third on the leaderboard, but withheld the submission on purpose as it was deemed unfair to other systems.

The grounding header, a fixed block of instructions, was created to improve performance on this benchmark without compromising fairness to other systems. The team is curious about the line between a grounding scaffold and coaching the grader and whether a fair test can determine the distinction.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Monday 24 August →