Five agents touched the answer. If it’s wrong, which one lied? Right now the honest answer is: nobody knows.

An LLM gives you an answer.

Where did it come from? No seriosuly, WHERE?

Not “which document” which transform. If five agents passed the claim down a chain before it reached you, which one introduced it? Which one paraphrased it into something slightly different? And when it turns out to be wrong, who is accountable?

We ship these pipelines to production and the answer is nobody. Every model output is an unprovenanced assertion. We’ve all got logs. Logs tell you what happened. They don’t tell you whether the transmission is worth trusting.

And multi-agent chains make this worse, not better. A claim that survives five hand-offs isn’t more reliable. It’s just more obscured.

The part I expected to be wrong

I assumed hallucination compounded with depth. Longer chain, more drift. It’s the intuitive story and it’s the one everyone tells.

So I built a leaderboard to measure it (issue #71 in the repo, all committed).

Hallucination rate on well-known facts: 0%. On post-cutoff facts: 44–56%.

And flat across chain depth.

The error doesn’t grow as it travels. It originates at hop 1 — at memory-generation, the moment the model reaches for something it doesn’t have. After that the relay is faithful. It propagates the error perfectly and adds nothing.

That changes where you put the instrumentation. You don’t need a drift detector. You need to know which hop the claim was born at, and whether that narrator had any business making it.

The discipline that already solved this

Fourteen centuries ago, Islamic scholars had the same problem with human claims and built an entire science around it. Hadith criticism. It runs three independent loops:

isnād — the chain. Who transmitted this to whom, in order.

rijāl — the narrators. A registry grading every transmitter’s reliability, by role, decayed for freshness.

matn — the content. Does this contradict what’s already established?

Three loops, deliberately separate. You can have a flawless chain carrying a false claim. You can have a true claim with a broken chain. The system refuses to collapse those into one number, because they’re not the same question.

ISNAD maps each loop onto LLM infrastructure.

Artifact: What it traces
Chain (isnād): every transform, in order, with input/output hashes
Registry (rijāl): narrator reliability, per-role, freshness-decayed
Grading: chain quality from the weakest narrator
Content criticism (matn): claim vs. evidence — plugs into your critic
Decision matrix: serve/caveat/review/quarantine
Corroboration: independent witnesses (tawātur), shared-error detection
Grounding: chain-scoped — what this chain actually retrieved
Content-madār fingerprints shared errors across supposedly-independent sources

Every claim routes through a 4×3 table — chain quality crossed with content verdict — and comes out as exactly one action. Serve. Serve with a caveat. Send to review. Quarantine.

Not a score. An action. You can put that in a pipeline.

It works as a library, a REST API, a CLI. It composes with LangChain, LlamaIndex, CrewAI, LangGraph and OpenTelemetry instead of replacing them. Nobody needs another framework asking them to rewrite their stack.

What it refuses to do

This is the part I care about most, and it’s the part most reliability tooling skips.

It grades WHO, not WHETHER. A ṣaḥīḥ chain means the transmission is well-attested. It never means the claim is true. That’s a different job and ISNAD doesn’t pretend to do it.

Ordinal, not numeric. Grades are ṣaḥīḥ > ḥasan > ḍaʿīf > mawḍūʿ. There is no “92.3% confidence,” because there’s no honest way to compute one. Anyone showing you two decimal places on trust is showing you a decoration.

It composes with critics. ISNAD is the transmission layer. It is deliberately not the content judge.

Measure, don’t assert. Every claim in the README traces to a committed, reproducible result, with the negative controls printed beside it.

Limits up front. Cold-start is worse per-role. Integrity is domain-scoped. Independence can’t be proven, only discounted. The critic is the coverage ceiling.

Those aren’t caveats I buried in a footnote. They’re in the README, above the fold.

The numbers

Content-madār fingerprinting: false-positive rate down to 0.375 from 0.750. Token recall 1.0. Near-miss 4/4.

Chain-grounding eval: grounded/ungrounded separation held under an NLI critic, with a degeneracy guard so it can’t cheat.

And the one I didn’t enjoy publishing: the critic is the ceiling. I ran a stronger cross-model critic — deepseek-v4-pro judging deepseek-flash — with the evidence sitting right there in context. It missed up to 2 of 4–5 hallucinations.

Your critic is not a substitute for adjudication. Mine isn’t either. The difference is that ISNAD says so in the docs instead of hoping you don’t check.

Who this is for

Teams shipping multi-agent pipelines who need per-hop accountability instead of a stack trace.

Medical, legal, financial — anyone who needs an audit trail, not a log file.

Researchers who want a preregistered, reproducible way to measure hallucination and provenance.

And anyone tired of “trust us, it’s fine” who’d rather have “here is who vouched for this, and how much they’re worth.”

pip install isnad          # Python ≥ 3.11, Apache-2.0
npm i isnad # JS verifier — verifies audit records, grading stays in Python

Repo: github.com/alizahidraja/isnad

One-command demo and the full benchmark harness are in there. So is the ARCHITECTURE doc, which is the honest map — the why, the limits, and the open problems I haven’t solved. Chain independence is still the big one.

Star it if you think LLMs deserve a chain of custody.

ISNAD grades the transmission, not the truth. The way reliability was done before machines, now done for them.