Urgent.News

What's breaking now, across thousands of outlets.

AI

Source-Aware Verification for MCP Agents: Why Fact-Checking Isn't Enough When Tools Lie About Provenance

Most fact-checking systems for LLM agents ask one question: is the claim supported by the evidence? They do not ask a second, equally important question: did the claim come from the source the agent cited? When an MCP agent pulls data from a search tool, a database query, a patient record API, and a policy document, then synthesizes an answer, a source-blind verifier will pass any claim that…

When an MCP agent aggregates data from various tools—search engines, databases, APIs, and documents—to generate responses, a standard fact-checker might deem the claim true if it appears somewhere among the retrieved evidence. However, ProvenanceGuard, a verification layer developed by Multiverse Computing, adds a crucial layer of scrutiny by examining not just the factual accuracy but also the provenance, or the origin, of the information.

This is particularly vital in domains like finance and healthcare where the source of the information can significantly impact the validity and trustworthiness of the answer.

The challenge stems from the Model Context Protocol (MCP), which allows agents to invoke multiple tools in a single turn. Each tool returns structured data, but the protocol lacks built-in mechanisms to sign responses, attest to data lineage, or assign confidence scores. Consequently, an agent might seamlessly combine facts from disparate sources—such as a web scraper, a PostgreSQL query, a Bloomberg API, and a PDF retrieval tool—without any explicit citation discipline.

This can lead to what is termed as "cross-source conflation," where a claim is accepted as valid even if it is based on information from an unverified or irrelevant source.

Addressing this, ProvenanceGuard intervenes after the agent has produced its answer but before the response reaches the user. It receives three key inputs: the agent's final answer, the set of outputs from each MCP tool, and the citations supplied by the agent—either explicitly or inferred. The system processes these inputs by first decomposing the answer into individual claims, then mapping each claim to the specific source from which it was derived.

Using a fine-grained natural language inference (NLI) model, ProvenanceGuard verifies whether each claim is indeed supported by the data from the cited source. If a claim is found to be supported by evidence from a different tool than the one it was cited from, the system flags this as cross-source conflation.

The system offers several actions in response to a flagged claim: it can block the entire answer, prompting the orchestrator to seek a more accurate response, or it can trigger a repair step. This repair step involves re-querying the correct tool to regenerate the claim, ensuring that the final response is based on verifiable and appropriately attributed information.

In implementing ProvenanceGuard, latency is a consideration. Running an NLI model on each claim can introduce a delay of 50-200ms per claim. Depending on the application, such as a customer support chatbot versus a high-frequency trading agent, this latency might be acceptable or prohibitive. To mitigate this, the system offers strategies like batch verification, where multiple claims are verified in parallel, selective verification focused on high-stakes sources, or caching entailment results for repeated tool outputs.

Additionally, while the paper presumes explicit citations (e.g., "According to the account record..."), many agents do not provide such explicit references. To address this, options include prompt engineering to enforce structured citations, heuristic matching to guess the likely source of each claim, or employing a second LLM to attribute claims to the most probable source.

Finally, when cross-source conflation is detected, ProvenanceGuard can take corrective measures. These include blocking the answer and returning an error, re-querying the correct tool for a verified statement, or downgrading the claim to a hedged form, such as "Some sources suggest..." However, the effectiveness of these repair mechanisms relies on the original tool's data being accurate.

If the cited tool does not contain the necessary information, the agent might either hallucinate or refuse to answer, highlighting the delicate balance required in this verification process.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

We quantized our AI judge. Here's exactly what broke.

Our production judge — a small 1.7B model with a LoRA adapter that grades other AI outputs as pass / fail / insufficient_evidence (88.5% accuracy, ECE 0.072) — is cheap to run.

  • Quantization implemented to reduce serving costs
  • Precision dropped to 98.28% and 94.16% with int8 and int4 formats
  • Four-rule deployment discipline established for production judges

Deterministic State Machines for Resilient Autonomous Agents

Deterministic State Machines for Resilient Autonomous Agents Autonomous multi-agent architectures routinely fail in production when relying on unconstrained large language model conversation loops.

  • Deterministic FSM replaces unconstrained LLM loops to prevent unpredictable behavior
  • AgentStateGraph governs deterministic control flow outside LLM reasoning core
  • Strict JSON validation and checkpointing enable state persistence and recovery

Headroom: How Context Compression Cuts Agent Token Costs by 60–95% Without Changing Answers

Production agents hit context limits fast. A coding agent that runs tests, reads logs, and pulls documentation can burn through 100k tokens in three turns.

  • Headroom compresses AI agent token usage by 60–95% without changing answers
  • Compression tool reduces coding agent token usage from 100,000 to below 5,000
  • Headroom maintains answer quality by preserving semantic anchors like error messages

More from Saturday 10 October →