
A 1,200-year-old method for grading who said what, rebuilt for AI systems that are wrong in public.
What it does, what it measured, and what it refuses to claim.
In 2025, Deloitte delivered a report to the Australian government on a contract worth just under AU$440,000. Readers found references to academic work that didn’t exist and a quote attributed to a judge that the judge never said. Deloitte confirmed a generative AI model had been used in the work, published a corrected version, and agreed to refund the final instalment. (AP News)
It isn’t an isolated story. One public database of court decisions involving AI-fabricated content had logged 2,022 cases by early September 2026, most of them in the United States.
The interesting question in each case isn’t whether the model hallucinated. Models hallucinate. The question nobody could answer quickly was the one any auditor, judge or client asks first:
Where did this come from, and who vouched for it on the way?
This piece is about the tool I built to answer that question, and about the 1,200-year-old discipline I took it from.
An answer is a relay race nobody filmed
A modern AI answer is rarely one model talking. A scraper pulls a page. A parser turns it into text. A retriever picks five chunks out of fifty thousand. A model writes a draft. A second agent summarises the draft. A tool formats it. Then a person reads one confident paragraph and forwards it.
Every one of those steps handled the claim. Every one could have bent it. And in most pipelines today, none of them leaves a record that says I touched this, here is what I did, and here is how much you should trust me.
We have observability tools, and they’re good. They log what happened: which call ran, how long it took, what it returned. That’s a trace. It isn’t a judgement. A trace tells you a retriever ran. It doesn’t tell you the retriever was pulling from a forum thread nobody had graded, and that this one weak hand is why the final answer is wrong.
Two things make this worse as systems grow
More hands don’t mean more truth. A claim that survives five hand-offs isn’t five times more reliable. It’s five times harder to trace. In my own earlier experiments, hallucinations didn’t pile up with chain depth; they were born at generation and then carried faithfully downstream, looking more official at every hop.
Agreement can be fake. Three agents agreeing feels like confirmation. But if all three read the same bad source, you don’t have three witnesses. You have one witness quoted three times.
So the problem isn’t only that models make things up. It’s that our systems can’t show their chain of custody, so nobody can say which link failed, or prove afterwards that the record wasn’t quietly edited.
Scholars solved a version of this a thousand years before GPUs
After the Prophet’s death, sayings attributed to him spread by word of mouth across a growing empire. Some were genuine. Many were not. The scholars who set out to separate the two faced exactly our problem: a statement arrives, sounding authoritative, and you need to know whether to rely on it.
Their answer was not to judge the statement by how convincing it sounded. Fluency was never evidence. Instead they built three separate disciplines, and refused to collapse them into one score.
The chain (isnād). Every narration came with its route: A heard it from B, who heard it from C, back to the source. No chain, no standing.
The narrators (‘ilm al-rijāl). Scholars compiled biographical records on tens of thousands of transmitters: when they lived, whom they could have met, how good their memory was, whether it declined with age, whether they were honest. Ibn Ḥajar al-ʿAsqalānī’s handbook Taqrīb al-Tahdhīb ranks narrators on twelve tiers, from impeccable down to rejected. An unknown narrator wasn’t given the benefit of the doubt. Unknown meant weak until shown otherwise.
The content (matn). Separately, scholars asked whether the text itself contradicted what was firmly established. A perfect chain doesn’t rescue an impossible claim.
Out of this came two rules that matter enormously for AI.
The first is the weakest link. A chain is only as strong as its least reliable narrator. Five excellent transmitters and one known fabricator do not average out to “pretty good”. The chain is graded ḍaʿīf, weak.
The second is about corroboration. A weak narration can be lifted by independent support from other chains. But scholars also tracked the madār: the point where supposedly different chains converge on a single narrator. If all your “independent” chains run through the same person, you don’t have corroboration. You have one source, echoed.
The tradition’s real insight wasn’t that transmitters were trustworthy. It assumed they weren’t, and built a system that made unreliability visible.
ISNAD: the same discipline, for machines
ISNAD is an open-source Python library (Apache-2.0) that applies those rules to AI pipelines. The mapping is close to literal.
Narrators become transmitters. Anything that handles a claim is a transmitter: a source corpus, a scraper, a retriever, a model, a tool, an agent. You register each one with a grade (reliable, acceptable, weak or ungraded) on two separate axes: integrity (does it alter or invent?) and precision (does it get details right?). Grades are set per domain and per role, because a model can be excellent at summarising and terrible at quoting.
Those grades are your priors, and ISNAD says so openly. It doesn’t pretend to know how good your scraper is. What it does is apply the discipline consistently once you’ve said what you believe. And it treats a model version bump as a new transmitter, because a new version hasn’t earned its predecessor’s record.
The chain rides with the claim. Every claim carries the ordered list of transmitters that handled it.
The weakest link caps the grade. As in hadith criticism, one ungraded or weak hand pulls the whole chain down. Unknown means weak by default. When I benchmarked this, the classical conservative rule beat my own more generous instinct, so the strict default stayed.
Corroboration is discounted when sources converge. If three agents agree but all three trace back to the same document, ISNAD counts that as one witness, not three. That’s the madār idea, and it’s the part of the method I find most underrated.
The output is a decision, not a confidence score. Every claim is routed to one of four outcomes: serve, caveat, review or quarantine. There’s no “92.3% confident”, because there’s no honest way to compute that number.
It fails closed. If a chain is ambiguous or can’t be graded, it goes to review. It is never silently passed.
Here’s the smallest version, from the README:
from isnad import Registry, grade
from isnad.types import NarratorGrade
reg = Registry()
reg.register("openstax", "physics", grade=NarratorGrade.RELIABLE)
reg.register("pdf-scraper", "physics", grade=NarratorGrade.UNGRADED)
reg.register("ingest-model", "physics", grade=NarratorGrade.ACCEPTABLE)
verdict = grade("p = mv", ["openstax", "pdf-scraper", "ingest-model"], reg, domain="physics")
print(verdict.why)
# claim 'p = mv' → chain DAIF (weakest: pdf-scraper, ungraded)
The textbook is excellent. The model is fine. But nobody ever graded the scraper, so the chain is ḍaʿīf, and ISNAD tells you exactly which hand is the problem.
It plugs into the stacks people already use: LangChain, where it’s listed as a community middleware integration in the official docs, plus LangGraph, CrewAI, LlamaIndex, MCP and OpenTelemetry.
The evidence, including the parts that don’t flatter it
A provenance tool that can’t show its own provenance is a joke. So every number below comes with its limits attached. A quick note on the unit: κ (Cohen’s kappa) measures agreement beyond chance. 1 is perfect, 0 is what you’d get by guessing. It is not a percentage.
Test one: the method on its home turf
I ran ISNAD’s weakest-link rule against 575,060 graded hadith chains, comparing its verdicts with a rule-based convention derived from Ibn Ḥajar’s twelve narrator tiers. The mappings were preregistered before the first result existed, so I couldn’t tune them afterwards.
- κ = 0.87 (3-way grading). 5-way: 0.8667. With the lenient setting: 0.761.
- Shuffled control: κ ≈ 0. Randomise the narrator grades and the signal disappears, which is what you want to see.
- The weakest number, published on purpose: at the level of individual narrator grades, agreement is only κ = 0.33. Human experts disagree a lot about individual transmitters. The rule is consistent; the inputs it inherits are contested.
That last line matters. κ = 0.87 means ISNAD applies a classical convention faithfully. It does not mean ISNAD knows which hadith are authentic. Agreement with a convention is not ground truth. (An earlier post of mine quoted 577,024 chains; that’s the raw corpus. 575,060 is the number actually graded.)
Test two: live AI answers
Hadith chains are friendly territory. The real question was whether the method transfers to messy LLM pipelines. I tested it on RAGTruth: 17,790 responses from 6 models, with human-labelled hallucinations.
- κ = 0.4345 (fail-closed) against the human labels, on a 1,800-response sample.
- 70.3% accuracy, against a 55.8% majority-class baseline.
- 96.0% of hallucinations flagged, at 60.4% precision. Recall isn’t precision — the critic also flags grounded answers as hallucinations 49.9% of the time. 96.8% of responses parsed; the 3.2% that didn’t are scored fail-closed, not excluded, and I say so every time I quote the number.
- It ranked four of the six models in the correct order; the bottom two near-tied models swap.
Here’s what those numbers don’t say. κ = 0.4345 is moderate agreement, not a solved problem. A 96.0% recall figure says nothing about false positives; recall isn’t precision, and I share the precision sheet on request rather than letting the big number stand alone. And a small share of responses couldn’t be graded at all. ISNAD routes those to review instead of guessing, which is the honest behaviour, but it also means a human still has work to do.
I’d rather you trust a modest number with its caveats than a perfect number without them.
A record you can hand over, and prove nobody changed
Grading a chain is half the job. The other half is being able to show the grading later, to someone who doesn’t trust you.
So every routing decision produces an AuditRecord: the claim, its full chain, each transmitter’s grade, the decision and the reason. Each record gets a SHA-256 hash of itself, a signature (HMAC or Ed25519), and a place in an append-only Merkle log. If anyone edits history, the hashes stop matching. The records export in an open format you can verify offline, with no dependency on me.
That makes the record tamper-detecting: it can show that it hasn’t been altered since it was written. I’m careful with that word, because the honest boundaries are the whole point:
- ISNAD does not tell you whether a claim is true. It tells you who handled it, how much you said you trust each of them, and what the system decided. A person still adjudicates.
- The grades are your priors. Grade a bad source “reliable” and the chain will inherit your mistake. The record will show that you made it, which is the point.
- “Signed” means signed with your key. It is not a third-party audit, and I won’t describe it as one until an audit exists.
- It is not a conformity assessment of anything. It produces evidence. Whether that evidence satisfies a regulation is a question for your counsel and your auditor.
The threat model is public too: sleeper narrators that behave well and then don’t, Sybil transmitters, poisoned grade imports, tail truncation, and the limits of what a forger can and can’t do. If your security team wants to attack it, the repo is open.
Why this matters now, stated carefully
There’s a lot of loose talk about the EU AI Act, so here are the dates as they actually stand after the Digital Omnibus entered into force on 27 July 2026:
- 2 August 2026: Article 50 transparency obligations already apply.
- 2 December 2027: obligations for stand-alone high-risk systems (Annex III: credit scoring, employment, education, essential services and more).
- 2 August 2028: high-risk AI embedded in regulated products (Annex I).
For high-risk systems, Article 12 requires the capability to log events automatically, and Article 19 requires providers to keep those logs for at least six months. (The ten-year figure people quote is Article 18, and it covers documentation, not logs.)
ISNAD doesn’t make anyone conform with the AI Act, and I’ll never say it does. What it produces is evidence of the kind those articles ask for: which components handled an output, what the system decided, and a record that can show it wasn’t edited afterwards. It can also feed the technical documentation a provider writes.
The practical point is simpler than the legal one. A log can only cover time after it exists. No library can produce a record for decisions made before it was installed. Teams scoping Annex III systems now have a year to find out which links in their chains they can’t account for, and that’s far better discovered in a test than in front of an auditor.
And outside Europe, the questions are arriving anyway. The UAE’s central bank published guidance in February 2026 expecting licensed financial institutions to run documented AI governance and offer people an explanation of AI decisions. Courts, clients and newsrooms are asking the same question without any regulation at all: where did this come from?
Grade the chain. Sign the claim.
If you run anything with more than one hand in it (a RAG pipeline, an agent swarm, a knowledge base that rewrites itself), try this on one real path today:
pip install isnad
Register the five or six things that touch an answer. Be honest about which ones you’ve never actually graded. Run a week of real questions through it and look at what lands in review and quarantine. That list is usually the most useful document a team has seen about its own system.
The core is Apache-2.0 and stays that way. If you’d rather have someone map and grade one of your pipelines with you, I run a two-week assessment that does exactly that and hands you the signed evidence at the end. Details are at isnadhq.com.
And if you think the method is wrong, break it in public. The benchmark, the controls, the threat model and the failures are all in the repo. The best improvements to ISNAD so far came from people who tried to knock it down.
The scholars who built hadith criticism didn’t trust their sources either. They just refused to let that distrust stay invisible. That’s all ISNAD is: the way reliability was done before machines, done for them.
Paper: arXiv:2607.24117 · Code: github.com/alizahidraja/isnad · Package: pypi.org/project/isnad · Site: isnadhq.com