Case studies

Case #0 — RAGTruth

Until a client case exists, here's the benchmark case — written the way every case study should be written: with the weak numbers printed at the same size as the strong ones.

Does the grounding critic transfer to a public benchmark?

RAGTruth is a public research dataset of 17,790 responses across 6 models, with human labels marking which answers contain hallucinated content. I ran ISNAD's grounding critic against a 1,800-response sample of it.

κ = 0.4345
Cohen's κ, fail-closed (1 = perfect, 0 = chance). The 3.2% that did not parse are scored fail-closed — not excluded.
96.0% recall
Hallucination recall — at 60.4% precision. Recall ≠ precision: the critic catches most hallucinations and also over-flags.
70.3% vs 55.8%
Accuracy against a 55.8% majority-class baseline.

What it shows

On a public benchmark the critic was not tuned on, it 4 of 6 in the right order; the bottom two near-tied models swap. A single check on a single answer is noisy; a track record built from many checks is sturdier — which is exactly why ISNAD grades transmitters over time instead of judging answers one by one.

What it does not show

This is a transfer result on a research dataset, not a client pipeline. It says nothing about your systems until the same discipline is run on them — which is the point of the Provenance Assessment.

Caveats, printed here on purpose

The critic catches 96.0% of hallucinations but flags grounded responses as hallucinations 49.9% of the time, and it over-flags between 1.3× and 3.4×. The headline κ is 0.4345 with unparseable responses scored fail-closed (1,743 of 1,800 parsed — 96.8%). None of these numbers appear in the arXiv paper; they live in the repository, in G1_CASE_STUDY.md, where anyone can rerun them.

Sources: https://github.com/alizahidraja/isnad · RAGTruth dataset (Cheng et al.). When a client case exists, with their written permission, it takes this page's place.

Want this run on your pipeline instead of a benchmark?