Case studies
Case #0 — RAGTruth
Until a client case exists, here's the benchmark case — written the way every case study should be written: with the weak numbers printed at the same size as the strong ones.
Does the grounding critic transfer to a public benchmark?
RAGTruth is a public research dataset of 17,790 responses across 6 models, with human labels marking which answers contain hallucinated content. I ran ISNAD's grounding critic against a 1,800-response sample of it.
What it shows
On a public benchmark the critic was not tuned on, it 4 of 6 in the right order; the bottom two near-tied models swap. A single check on a single answer is noisy; a track record built from many checks is sturdier — which is exactly why ISNAD grades transmitters over time instead of judging answers one by one.
What it does not show
This is a transfer result on a research dataset, not a client pipeline. It says nothing about your systems until the same discipline is run on them — which is the point of the Provenance Assessment.
Caveats, printed here on purpose
The critic catches 96.0% of hallucinations but flags grounded responses as hallucinations 49.9% of the time, and it over-flags between 1.3× and 3.4×. The headline κ is 0.4345 with unparseable responses scored fail-closed (1,743 of 1,800 parsed — 96.8%). None of these numbers appear in the arXiv paper; they live in the repository, in G1_CASE_STUDY.md, where anyone can rerun them.
Sources: https://github.com/alizahidraja/isnad · RAGTruth dataset (Cheng et al.). When a client case exists, with their written permission, it takes this page's place.