Urgent.News

What's breaking now, across thousands of outlets.

AI

I've Built RAG Infrastructure Several Times. Last Week Was the First Time I Actually Benchmarked It.

I have a confession, and I suspect I'm not alone in it: I've built RAG infrastructure multiple times, and until last week I had never benchmarked any of it. Unit tests, sure. Integration tests, sure. Everything green, every pipeline connected. But if you'd asked me "is the retrieval actually good?", the honest answer was a shrug with a deployment attached. For context: I'm building Langhuan , an…

I have a confession: I've built RAG infrastructure multiple times before, yet I never benchmarked it until last week. Unit tests, integration tests, everything appeared green, and pipelines were connected. However, when asked if the retrieval was actually good, I shrugged and said, "Sure, deployment attached."

To address this, I started auditing the retrieval decisions I had shipped. Chunking contracts had undergone three revisions, and the fusion of vector and keyword search used RRF. A reranker was placed on top. However, none of these decisions were backed by numbers. My approach was based on architectural taste, and taste doesn't fail loudly.

The expensive part of an evaluation isn't the harness; it's the labeled data. To label nothing, I deterministically sampled 200 real queries from MIRACL-zh, a Chinese Wikipedia corpus with human-annotated passage relevance, under the Apache-2.0 license. I ran four configurations: vector only, FTS only, hybrid, and hybrid + rerank. Metrics included recall@10, MRR@10, and nDCG@10 against fixed qrels. Determinism was crucial; the same fingerprint, model versions, and code must produce identical metrics across runs.

To validate the harness, I first ran a smoke test with a mock embedding. Scores closely matched the random baseline, indicating the harness wasn't biased. However, the harness also needed to spin up a real standalone instance, create a knowledge base, and write to it. Surprisingly, this path consistently returned a 500 error. The issue stemmed from a missing deferred check on SQLite, causing the forward foreign key to fail.

Next, I discovered that the vector search had never worked in the production binary. The vec extension was missing from the build. This oversight went unnoticed by existing tests, as they mocked the database and ran against Postgres. When the real model was introduced, the results table displayed a row of 0.0000s, indicating zero recall in the FTS channel.

The hybrid search, however, scored 0.9799, identical to the vector-only configuration. This discovery revealed that my hybrid search had been running as a plain vector search with additional steps, unnoticed for an extended period.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Why LLM reasoning isn't enough for medical scheduling math

I’ve seen plenty of people try to make Claude or GPT-4 act like a specialized scheduler. They prompt it heavily: "You are a precise medical assistant.

  • LLMs excel at simple scheduling but struggle with complex constraints
  • Probabilistic nature of LLMs leads to hallucinations in deterministic scheduling
  • Injection Day Alignment MCP server offers structured tools for medical scheduling

Passing Once Isn't Reliable — This Week's Agent Engineering Puts the Harness Before the Model

This digest covers AI agent developments from 2026-08-18 to 2026-08-25: orchestration patterns, tool/function calling, memory, planning loops, multi-agent coordination, and agent evaluation.

  • AgentWeave reduces tool exposure by 70% and latency by 51%
  • Pass@1 in Thinkingbox drops from 65.36% to 25.25% despite initial success
  • AutoSaddler optimizes agent harness, gains 9-10 points across benchmarks

More from Tuesday 25 August →