Retrieval for law and tax
- Question
- How does a retrieval system know that a source which looks relevant does not apply?
- Status
- Working paper with a proposed architecture and four falsifiable predictions. Evaluation is the next step.
Similarity scores alone do not establish jurisdiction, version validity or conceptual applicability. A treaty article from the right country but the wrong year, or a provision from a neighbouring jurisdiction in near-identical words, scores as highly as the right source, and a generator then answers from it with full confidence.
Approach
- Typed negative evidence: a source can fail for three distinct reasons, and each needs a different response.
- The working paper proposes calibrated relevance probabilities as the output of a merge scorer, with the type and severity of uncertainty selecting the response: assert, hedge, flag a version boundary, or acknowledge a gap.
- Detection from metadata and structure, not from a second language model, so that the check is deterministic and auditable.
| Failure | The source is | The system should |
|---|---|---|
| Scope | on topic, for another jurisdiction or entity | exclude it |
| Sibling | about a related but different concept | disambiguate |
| Temporal | correct once, now superseded | resolve the version |
Related work in progress
- A bitemporal record store that answers "what was known on this date", tested on regulatory filings. The first pre-set hypothesis came back void, which is recorded as such. It also found a genuine look-ahead leak: filing dates diverge from acceptance dates, in some cases backwards by months.
Paper: What doesn't match matters more, with DOI.