I Measured My RAG Pipeline Honestly. It Was 40x Slower Than I Thought.
A few days ago I published the architecture behind Vicquant’s RAG Vault; a 9-stage retrieval pipeline built to ground financial AI answers in actual source documents, with strict citations, so a user asking “what’s my 401(k) contribution limit” gets an answer tied to a real document, not a model’s confident guess. The quality evaluation backed it up: +5.2% answer relevance, +16% context…
A few days ago, the author published the architecture behind Vicquant’s RAG Vault, a complex retrieval pipeline designed to ground financial AI answers in actual source documents. The quality evaluation confirmed the pipeline's effectiveness, with improved answer relevance, context precision, and context recall compared to a simpler version.
However, the author lacked a clear understanding of how fast the pipeline actually was. To find this out, they conducted a live benchmark using the same 25-question golden evaluation set, running the real pipeline against the actual API three times. This revealed a mean latency of 14.81 seconds, with a p95 of 21.28 seconds. The author then identified and addressed several bottlenecks in the pipeline.
The first major fix was routing, which classified document-grounded questions into the full pipeline and other questions into a faster direct-chat path. This routing process added only 0.09 milliseconds of overhead while saving significant time. Next, the author improved concurrency by running HyDE generation and multi-query expansion concurrently using asyncio.gather, which cut roughly 20% off the document-grounded path.
Streaming was also introduced, allowing the fast path to render its first token in under a second. After implementing these changes, the mean latency dropped to 6.24 seconds, and the median latency reduced to 3.61 seconds. However, despite these improvements, the author still noticed inconsistent response times. This inconsistency stemmed from OpenRouter's unpredictable routing of model requests to different inference providers, with one provider being about 11 times slower than others.
To resolve this, the author pinned the most efficient providers at the beginning of the request queue, resulting in a 97.5% reduction in mean latency from the initial measurement. The main takeaway from this experience is that a benchmark must accurately measure the user experience, and "random" variance is often a result of underlying infrastructure factors that can be identified and optimized.
Ultimately, these optimizations enhanced speed without compromising the pipeline's quality, demonstrating that speed and correctness are not inherently in tension.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
