Corpus token coverage
Blocks a response when its overall unsupported content-token share exceeds a train-calibrated threshold.
TP 721 · FP 645 · TN 1136 · FN 173
Measured report · 01
We tested three reproducible local gates on RAGTruth. All caught most annotated answers. None was precise enough to approve a production release alone.
2,675
held-out responses
894
annotated unsupported
82.5%
highest recall
64.9%
highest F1
The hybrid gate had the best response-level F1 at 64.9%, but it still missed 162 of 894 annotated test responses and blocked 631 supported responses.
That is the useful conclusion: lexical evidence can be a cheap pre-filter, not a production acceptance oracle. A consequential workflow still needs task-specific claims, adjudication policy and a held-out release set.
Measured result
Thresholds were calibrated on the official train split by maximum F1, then frozen for the official test split.
| Gate | Threshold | Precision | Recall | F1 | False-negative rate | Mean feature time | Provider cost |
|---|---|---|---|---|---|---|---|
| Corpus token coverage | 0.34 | 52.8% | 80.7% | 63.8% | 19.4% | 0.0229 ms | $0.00 |
| Weakest-sentence support | 0.58 | 51.8% | 82.5% | 63.6% | 17.4% | 0.0229 ms | $0.00 |
| Hybrid claim gate | 0.44 | 53.7% | 81.9% | 64.9% | 18.1% | 0.0229 ms | $0.00 |
Latency is the observed local feature-extraction time divided by eligible responses on the execution machine; it is not a cross-hardware performance claim.
Three gates
Blocks a response when its overall unsupported content-token share exceeds a train-calibrated threshold.
TP 721 · FP 645 · TN 1136 · FN 173
Blocks a response when any sentence of five or more content tokens falls below the train-calibrated support threshold.
TP 738 · FP 687 · TN 1094 · FN 156
Combines overall support, weakest-sentence support, and source-missing numeric claims in a fixed weighted score.
TP 732 · FP 631 · TN 1150 · FN 162
Failure analysis
It flagged 645 supported responses because legitimate paraphrases and ordinary connective language are absent from the context.
It achieved the highest recall, 82.5%, while sending 687 supported responses to review.
Adding numeric mismatch improved F1, but 162 annotated responses still passed. Source vocabulary can be rearranged into an unsupported claim.
Frozen protocol
Download evidence
No email gate. The CSV contains response IDs, manual labels and each gate’s decision; source text remains in the licensed upstream corpus.
Planning a RAG system? Use the project estimator to define the evidence, permission and release work around it.
Estimate a Generative AI build