Urgent.News

What's breaking now, across thousands of outlets.

AI

I Measured My RAG Pipeline Honestly. It Was 40x Slower Than I Thought.

A few days ago I published the architecture behind Vicquant’s RAG Vault; a 9-stage retrieval pipeline built to ground financial AI answers in actual source documents, with strict citations, so a user asking “what’s my 401(k) contribution limit” gets an answer tied to a real document, not a model’s confident guess. The quality evaluation backed it up: +5.2% answer relevance, +16% context…

A few days ago, the author published the architecture behind Vicquant’s RAG Vault, a complex retrieval pipeline designed to ground financial AI answers in actual source documents. The quality evaluation confirmed the pipeline's effectiveness, with improved answer relevance, context precision, and context recall compared to a simpler version.

However, the author lacked a clear understanding of how fast the pipeline actually was. To find this out, they conducted a live benchmark using the same 25-question golden evaluation set, running the real pipeline against the actual API three times. This revealed a mean latency of 14.81 seconds, with a p95 of 21.28 seconds. The author then identified and addressed several bottlenecks in the pipeline.

The first major fix was routing, which classified document-grounded questions into the full pipeline and other questions into a faster direct-chat path. This routing process added only 0.09 milliseconds of overhead while saving significant time. Next, the author improved concurrency by running HyDE generation and multi-query expansion concurrently using asyncio.gather, which cut roughly 20% off the document-grounded path.

Streaming was also introduced, allowing the fast path to render its first token in under a second. After implementing these changes, the mean latency dropped to 6.24 seconds, and the median latency reduced to 3.61 seconds. However, despite these improvements, the author still noticed inconsistent response times. This inconsistency stemmed from OpenRouter's unpredictable routing of model requests to different inference providers, with one provider being about 11 times slower than others.

To resolve this, the author pinned the most efficient providers at the beginning of the request queue, resulting in a 97.5% reduction in mean latency from the initial measurement. The main takeaway from this experience is that a benchmark must accurately measure the user experience, and "random" variance is often a result of underlying infrastructure factors that can be identified and optimized.

Ultimately, these optimizations enhanced speed without compromising the pipeline's quality, demonstrating that speed and correctness are not inherently in tension.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Noticed: a quiet photo diary with captions from a local AI

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What I Built Noticed is a quiet photo diary that runs entirely on your own computer.

  • Noted is a local AI-powered photo diary
  • Gemma 3 model generates captions locally
  • Photos and captions never leave user's computer

Touch Grass Birder: On-Device Bird ID in the Browser with Open-Weight AI

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass * What I Built Touch Grass Birder is an offline, on-device bird call identifier that runs entirely in your…

  • Touch Grass Birder identifies bird species via audio calls
  • Uses open-source AI (BirdNET v2.4) on user's device
  • No server needed; runs locally for privacy and zero cost

My app accused the rain of being a car.

How do you prove to someone that you actually touched grass? I guess the bigger question is out of everything going on how is "if i touched grass" still the biggest argument we are having? No more.

  • App "Proof of Grass" accuses rain of being a car
  • Offline app uses Perch 2.0 model to analyze sounds
  • Rain identified as human-made, scoring below pass mark

More from Saturday 10 October →