Skip to content

Case study · G1

Does source-grounding transfer to hallucination detection?

Can "grade who transmitted the claim" detect when an LLM fabricates content not present in its source? Two signals, honest results — on RAGTruth (17,790 responses, 6 models, span-level labels).

v1 — bi-encoder cosine (cheap, local)

`all-MiniLM-L6-v2` similarity between each sentence and its source; weakest-link over sentences.

ThresholdCohen's κAccuracy
0.3 (best)0.21559.7%
0.5 (preregistered)0.09951.3%

Verdict: too weak — semantic similarity is not entailment, and long sources get truncated.

v2 — LLM grounding critic (DeepSeek V4 Flash)

A strict grounding judge: "is this response fully grounded in its source?"

Cohen's κ0.4345 (fail-closed, 1,800) · 0.468 (parsed-only, 1,743)
Accuracy70.3% (majority-class baseline 55.8%)
Precision (halluc.)60.4%
Recall (halluc.)96.0%
F1 (halluc.)74.1%

Known limitation (stated, not hidden): 57 responses (3.2%) scored fail-closed as hallucinated, and a 49.9% false-positive rate on grounded responses (501/1,005) sits beside the 96.0% recall — the cost is review, not a wrong serve.

The narrator-grading signal — rank-orders models, but over-flags

The critic rank-orders the six models roughly right (4/6 exact), but predicted rates run 1.3–3.4× the true rates.

ModelTruthLLM-critic predicted
gpt-4-061312.3%39.7%
gpt-3.5-turbo-061316.0%45.3%
llama-2-70b-chat46.0%76.3%
llama-2-13b-chat57.0%83.7%
llama-2-7b-chat66.3%90.0%
mistral-7B-instruct67.3%86.3%

Overall verdict (G1 gate)

The LLM grounding critic transfers at κ = 0.4345 — a moderate signal with a real false-positive cost. It does not clear the κ ≥ 0.8 company-path bar yet, but it is a usable high-recall, low-precision screening signal and an honest first case study. Next lever: reduce the false-positive rate on grounded responses — the 49.9% FP cost is the real blocker, not parse loss.