Urgent.News

What's breaking now, across thousands of outlets.

AI

RAG vs. Direct Context: I Tested Both on Real Documents, Here's What Broke

A hands-on test of BGE-M3 + Qwen3 (RAG vs. direct-context answering) on a real research paper and a full-length book including a retrieval bug hiding in a footnote, and one surprisingly good model behavior. I wanted to answer a simple question: when you feed a document to an AI model, is it actually reading it or just pattern-matching to whatever text happens to look similar to your question? So…

I conducted a hands-on test comparing BGE-M3 + Qwen3 (RAG vs. direct-context answering) on real documents, including a research paper and a book. The goal was to determine if AI models are actually reading the document or merely pattern-matching to text that resembles the question. I built an open-source pipeline to generate separate answers for each approach.

For the research paper, both RAG and direct answers agreed on the topic of the paper, which presented a transformer-based bidirectional machine translation system for the legal domain. The RAG approach retrieved chunks from the document, while the direct answer read the raw text. The retrieval process was straightforward, using fixed-size chunking and plain cosine similarity without any fancy tricks.

However, when testing a full-length book, RAG failed to provide an accurate answer. It mistakenly identified a footnote in the dedication as the paper's main contribution, which was almost word-for-word a title of a blog post. This issue stemmed from the model not filtering out the front matter, such as dedications, acknowledgments, and footnotes, before processing the text. In contrast, the direct answer correctly identified the document as a book review and promotional content for a specific work.

On a more positive note, when asked a more specific question about the book, RAG performed exceptionally well. It accurately stated that the retrieved context did not include a relevant answer to the question about the main contribution of the paper. This demonstrated that RAG can work effectively when the retrieved context genuinely lacks the necessary information, as opposed to providing incorrect or misleading answers.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Why RAG on legal text keeps hallucinating dates - and what actually fixed it

A couple of weeks ago I dropped the CRA text (the EU's cybersecurity regulation for IoT devices) into ChatGPT and asked when the main requirements actually kick in.

  • ChatGPT incorrectly stated CRA regulation entered force in 2024, 3 years off.
  • Dispersed dates in EU regulation PDF caused AI confusion.
  • Platanor rebuilt repository with proper file structure to fix RAG hallucinations.

Data center backlash echoes fossil-fuel politics

Willie Nelson has become an unlikely barometer of America's biggest infrastructure fights. He once protested the Keystone XL pipeline and fracking. Today, he's fighting data centers. Why it matters: The American icon's latest cause underscores how the politics surrounding AI infrastructure are beginning to resemble the last decade's…

AI scrambles the political map

The search for a winning message on AI is pushing candidates and lawmakers into unexpected political territory. Why it matters: With the midterms approaching, AI is creating alliances across party lines while opening fissures within them. Here are three ways AI is scrambling the political map. 1.

More from Friday 14 August →