When does an agent beat RAG? We benchmarked three pipelines on a TigerGraph knowledge graph
Agentic GraphRAG Hackathon by TigerGraph, Round 1. Code: https://github.com/anirudh12032008/agentic-graphrag-tigergraph Ask a RAG system "How many biathlon events at the 2018 Winter Olympics had more than 73 competitors?" and it will give you a confident number. It will also be wrong. The answer depends on 8 to 43 documents, and a top-8 vector search can't see all of them. We built three…
When faced with certain types of knowledge graph queries, traditional Retrieval-Augmented Generation (RAG) systems may struggle to provide accurate answers, while agent-based GraphRAG systems can surpass RAG in both accuracy and efficiency. Researchers from TigerGraph conducted a benchmark using three pipelines on the same Claude-sonnet-5 language model, embeddings, and TigerGraph Savanna instance.
The first pipeline, RAG, employed a top-8 vector search followed by a single LLM call, resulting in a 48% accuracy on 100 public questions. The second pipeline, GraphRAG, used a top-6 vector search, potentially up to three seed events, followed by a fixed graph expansion before another LLM call, achieving 69% accuracy. The third pipeline, Agentic GraphRAG, was an LLM orchestrator that performed tool-based steps such as entity linking, GSQL aggregation, graph traversal, venue and date lookup, vector search, and document reading, ultimately scoring 99% accuracy.
GraphRAG performed better than RAG on aggregation and superlative questions, while Agentic GraphRAG outperformed both on multi-hop questions. The agent utilized GSQL aggregation and graph traversal to solve specific question types, achieving 100% accuracy for lookups, 67% for aggregations, and 90% for superlatives. In contrast, GraphRAG only achieved 80% accuracy on aggregation questions.
The agent also outperformed RAG in multi-hop questions, scoring 100% against RAG's 21%. Agentic GraphRAG demonstrated an ability to flag real ambiguity, such as with multiple finals at the same venue on different dates, and reported all candidates with citations instead of guessing. It even caught a data tie in fencing competitions, reporting both results with citations.
Compared to RAG, Agentic GraphRAG required 1.24 times more tokens but delivered 2.06 times the accuracy, making it 40% cheaper per correct answer. Additionally, the agent ensures that all answers are backed by citations, preventing it from relying on memory alone. All components, including the graph schema, GSQL queries, pipelines, benchmarking tools, and a Streamlit dashboard, are open source and readily available for replication.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.