Urgent.News

What's breaking now, across thousands of outlets.

Tech

You Might Not Need a Vector Database

A RAG pipeline is a lot of parts: a chunker, an embedding model, a vector database, a retriever, usually a reranker, and an eval harness to tell you when retrieval quietly got worse. Cache-augmented generation (CAG) deletes all of it. You put the entire knowledge base in the prompt, cache it at the provider, and ask your question. No retrieval step, so no retrieval mistakes. That sounds like a…

A RAG pipeline typically includes several components: a chunker, an embedding model, a vector database, a retriever, a reranker, and an evaluation system. However, the "Cache-Augmented Generation" (CAG) approach simplifies this by loading the entire knowledge base into the model's context and answering queries directly from that context without a separate retrieval step. This eliminates retrieval errors altogether.

In simpler terms, CAG involves putting the complete corpus at the beginning of the prompt, followed by a cache breakpoint, and then the user's question after the breakpoint. This allows the model to answer questions using the cached information without having to retrieve data from a separate vector database. The primary constraint for CAG is the cache's time-to-live (TTL), which is usually set to a short duration, like 5 minutes, and not the model's context window size.

CAG is not a framework to install but rather a prompt layout strategy. The corpus is placed at the beginning of the prompt, followed by the user's question after a cache breakpoint. This setup ensures that once the corpus is cached, subsequent queries can be answered without re-reading the corpus, reducing costs and eliminating retrieval mistakes. CAG can be more cost-effective than traditional RAG methods, especially for smaller corpora or when the corpus changes infrequently.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

Ollama Connection Refused? The 60-Second Triage

Originally published on mrsaynothing.dev This shipped today on mrsaynothing.dev — the agent-run site's daily post. Receipts live there first.

  • Ollama reports "Connection refused" due to server not listening on port 11434.
  • Verify Ollama service status with systemctl status ollama or ollama list.
  • Wrong port/host or Docker container issues cause connection errors.

I spent 6 months building a real operating system for web apps

For the last six months or so I've been building something a bit unusual: a real operating system for web apps. It's called PhreshOS.

  • PhreshOS is an open-source operating system for web apps
  • Users access via browser, no HTTP/WS setup needed
  • Supports AI agents interacting with system directly

More from Wednesday 30 September →