Urgent.News

What's breaking now, across thousands of outlets.

AI

When do AI agents actually matter? Benchmarking RAG vs GraphRAG vs Agentic GraphRAG on TigerGraph

Everyone's building agents right now. But a better question than "can an agent do this?" is "when does an agent actually beat something simpler?" For the TigerGraph Agentic GraphRAG Hackathon, I built three question-answering pipelines side by side to find out. The setup The dataset was ~2,900 Wikipedia articles about Olympic events, plus 100 evaluation questions with answers and 50 hidden ones.…

Three question-answering pipelines were built side by side for the TigerGraph Agentic GraphRAG Hackathon to determine when AI agents outperform simpler methods. The dataset consisted of nearly 2,900 Wikipedia articles about Olympic events, 100 evaluation questions with answers, and 50 hidden questions. The questions were categorized into five types: simple lookups, multi-hop, temporal, aggregations, and superlatives.

A plain text-retrieval system struggled with distractor articles, while a graph built solely from Olympic event records avoided them. The core idea was that the LLM planned the process while the graph computed the answers. Each event's Wikipedia infobox was parsed into a structured Event vertex, and 2,187 of these were loaded into TigerGraph Savanna.

Three pipelines were implemented: Retrieval-Augmented Generation (RAG), GraphRAG, and Agentic GraphRAG. RAG retrieved the top 5 documents by similarity and extracted the answer from them. GraphRAG executed a single graph query and returned the result. Agentic GraphRAG employed an orchestrator loop that planned, queried the graph, judged the sufficiency of evidence, self-corrected if necessary, verified against a second source, and stopped when confident.

Results showed that RAG scored 18% accuracy with an estimated 1,573 tokens per query. GraphRAG achieved 92% accuracy at a fifth of RAG's token cost. Agentic GraphRAG reached 100% accuracy but at the cost of 1,295 tokens per query. RAG struggled with aggregation and superlative questions, while GraphRAG excelled in multi-hop questions, achieving 100% accuracy. The takeaway is that agents are not universally superior; they demonstrate their value when dealing with ambiguous information requiring verification.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Saturday 3 October →