For Tool-Using Agents, a True Claim Can Still Have the Wrong Provenance

ProvenanceGuard evaluates whether each claim is supported by the source an agent attributes it to. The research also offers a practical blueprint for testing source-aware verification without confusing a conservative gate with proof of correctness.

Source artwork for Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
Source artwork · Hugging Face / credited contributors ↗
THE SHORT VERSION

Preserve source identity throughout an agent trace, evaluate support and attribution separately, and choose block-and-repair thresholds according to the cost of each failure.

A September 29 research article from Multiverse Computing introduces ProvenanceGuard, a post-generation verifier for agents that use multiple Model Context Protocol sources. Instead of pooling every tool result into one context, the method decomposes an answer into claims, routes each claim to a likely source, checks support and attribution, and can block or repair the answer. In a held-out medical evaluation, the authors report that it caught 138 of 139 claims experts said should not pass, while also holding 67 supported claims for review. On claims with an identifiable source, its source selection was correct about 86% of the time. A harder similar-source test reduced exact-source accuracy to 50.3%, an important limit alongside the headline results.

The central engineering lesson is that factuality and provenance are related but distinct. An answer can contain a true statement and still mislead users about where it came from. That distinction matters whenever one source has different authority, freshness or access controls from another.

Model evidence as labeled records

A source-aware verifier needs more than the final prose and a pile of retrieved text. Each tool response should retain a stable source identifier, tool name, retrieval time and, where practical, a location within the returned material. Those labels must survive transformations such as filtering, summarization and reranking.

Consider an account assistant with access to a customer's record and a general refund policy. The statement that a refund window exists may be supported by the policy, but phrasing it as an attribute of the customer's account changes its meaning. A useful trace therefore records both the content and the boundary between those sources. Concatenating them before verification destroys the very distinction the reviewer needs.

Source identity also needs an explicit level of granularity. A website domain may be too broad; a particular document version may be appropriate. A database name may be too broad when rows have different owners or timestamps. Teams should define what counts as a source according to the decisions the answer can influence.

Test three failures separately

Evaluation should distinguish unsupported content, incorrect attribution and unsuccessful routing. These failures can produce the same blocked verdict but require different fixes.

An unsupported claim has no adequate evidence in the captured trace. The safest response may be to remove it, qualify it or obtain more evidence. An attribution error occurs when evidence exists, but the answer names or implies the wrong source. That may call for corrected wording rather than deletion. A routing error occurs when the verifier checks the wrong source even though the answer and evidence could be valid. Improving the verifier's retrieval stage is then more relevant than changing the generator.

A compact test set can expose the differences. Start with answers whose claims have human-labeled supporting source IDs. Add controlled variants that preserve the claim while swapping the named source, and variants that change a critical value while retaining the attribution. Include several sources with overlapping language, because obvious mismatches make source routing look easier than it will be in production.

Report per-claim recall for claims that must be blocked, false-block rates for supported claims, and exact source-selection accuracy. An answer-level pass rate alone hides whether the system is useful or merely cautious.

Calibrate the gate to the consequence

A conservative verifier trades fewer unsafe passes for more valid answers sent to review or fallback. That can be appropriate for a clinical summary or a financial workflow, but frustrating for a low-stakes research assistant. The threshold should follow the consequence of a mistaken answer, not a universal idea of acceptable accuracy.

Define what happens after a block before deploying the gate. Options include showing a narrowly scoped fallback, rewriting only the failed claim, requesting another tool call or escalating to a reviewer. A repair should be verified again against the same labeled evidence. Otherwise the system may replace a visible attribution error with a smoother unsupported claim.

Latency belongs in this decision too. A verifier that runs after generation adds another critical path. Measure decomposition, routing and support checks independently, then decide which can be batched or cached without reusing a verdict across changed evidence. Offline review can tolerate a different budget from an interactive assistant.

Preserve auditability without overselling it

Per-claim verdicts are valuable because a reviewer can inspect which evidence was checked. They are not proof that the source itself is correct, current or authorized for the task. Source-aware verification answers a narrower question: whether the response is grounded in the identified evidence and attributes that evidence consistently.

Keep the original answer, claim segmentation, candidate sources, chosen source, support score, attribution decision, policy threshold and final action together. Version the verifier and its configuration. This makes later error analysis possible when a model update changes claim boundaries or routing behavior.

The strongest deployment pattern is layered: validate tool permissions before retrieval, preserve provenance through generation, verify claims after generation, and retain a clear fallback. ProvenanceGuard provides evidence that the post-generation layer can catch source mistakes that pooled-context checks miss. Its similar-source results also show why teams should test their own source landscape rather than treating a benchmark score as a deployment guarantee.

Source: Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents ↗. How we write

← Back to all articles