Donald Murre

What doesn't match matters more

Abstract

Retrieval-Augmented Generation (RAG) systems retrieve documents by positive relevance (how similar a document is to a query) and generate answers from whatever scores highest. This paradigm has no mechanism for representing what a document does not answer. We argue that this omission is the root cause of the most dangerous RAG failures: confident wrong answers in high-stakes domains.

We present evidence from legal, medical, and financial retrieval that positive-only scoring fails in three systematic ways: (1) it cannot distinguish topically similar but jurisdictionally inapplicable documents (scope failure); (2) it cannot distinguish related but categorically different concepts (sibling failure); and (3) it cannot distinguish current from superseded versions of a document (temporal failure). Each failure mode produces false positives that are invisible to the generation layer.

We derive three properties that any retrieval system adequate for high-stakes domains must satisfy: (1) typed negative evidence representation, capturing specific ways a document fails to answer a query; (2) calibrated retrieval confidence, producing P(relevant) rather than ranking scores; and (3) confidence-gated generation, selecting response strategies based on the type and severity of uncertainty. We define calibrated retrieval formally and propose a three-type taxonomy of negative evidence: scope, sibling, and temporal negatives.

As proof by construction, we present CRANE (Calibrated Retrieval with Adversarial Negative Evidence), a functional architecture satisfying all three properties. We state four falsifiable predictions and invite empirical evaluation.

Three ways a document fails to apply

TypeThe document is…Example from the paper
Scopeon the right topic, for the wrong entity or jurisdictionSwiss domestic tax law retrieved for a question on the Swiss–Germany double taxation treaty
Siblingabout a related but different conceptpermanent establishment retrieved for a question about residency
Temporalcorrect once, but supersededthe pre-2024 treaty text retrieved for a question about post-2024 remote workers

The pre-2024 and post-2024 versions of the treaty share roughly 95% of their text. An embedding of the text cannot tell them apart; the difference sits in metadata.

Four predictions that could prove it wrong

  1. A retrieval system with typed negative evidence produces fewer false-positive retrievals on cross-jurisdiction legal queries than the same system with positive-only scoring, measured by Document-Level Retrieval Mismatch.
  2. A merge scorer that produces calibrated P(relevant) supports meaningful confidence thresholds: an expected calibration error below 0.10, which positive-only scorers do not reach.
  3. A generation step gated on confidence profiles gives fewer assertive answers where the evidence is weak or contested than one that receives a flat ranked list.
  4. Typed temporal negatives reduce wrong-version retrieval on versioned legal and regulatory corpora.

If a positive-only system matches CRANE on all four, the argument is wrong. I would welcome that result.

Claude was used for light editing and rephrasing. The ideas, the argument and the sources are my own.

Cite

Murre, D. (2026). What doesn’t match matters more: CRANE, Calibrated Retrieval with Adversarial Negative Evidence. Working paper, 27 pages. Zenodo. https://doi.org/10.5281/zenodo.23134928

@misc{murre2026crane,
  author = {Murre, Donald},
  orcid  = {0009-0008-7246-0053},
  title  = {What doesn't match matters more: {CRANE}, Calibrated Retrieval with Adversarial Negative Evidence},
  year   = {2026},
  month  = apr,
  note   = {Working paper, 27 pages},
  doi    = {10.5281/zenodo.23134928},
  url    = {https://doi.org/10.5281/zenodo.23134928}
}