Research · Open Source
إسناد

ISNAD

Claim-level trust for multi-agent AI systems.

Every claim carries its chain. Every transmitter is graded. Adapted from 1,200 years of hadith transmission science.

arXiv:2607.24117 · cs.AI / cs.MA · 25 pages · pip install isnad v2.24.0 · Apache-2.0
PyPI version PyPI downloads per month npm version GitHub stars License arXiv
In Production

ISNAD in production — audit export#

The reference implementation now emits a tamper-evident AuditRecord for every graded claim — the evidence artifact auditors and regulators will ask for. ISNAD produces evidence artifacts; it does not confer conformity with any regulation.

isnad export --claim <claim-id> --format json --verify

Truncated example record (illustrative):

{
  "record_id": "audit-…",
  "record_version": "1",
  "claim_id": "c1",
  "claim_text": "p = mv",
  "final_grade": "hasan",
  "chain": [
    { "narrator_id": "src", "narrator_type": "dataset", "grade": "reliable" },
    { "narrator_id": "m", "narrator_type": "model", "grade": "acceptable" }
  ],
  "weakest_link": { "narrator_id": "m", "why": "lowest grade" },
  "integrity": { "record_hash": "sha256:…", "hash_algorithm": "sha256" }
}

The record hash is SHA-256 over RFC 8785-canonical JSON. Verify it with isnad verify --record record.json. The full field mapping — which AuditRecord field supports which requirement under the EU AI Act, ISO/IEC 42001, NIST AI RMF, and SDAIA — is at docs/evidence-mapping.md.

What ISNAD deliberately does not provide: no conformity statement, no risk classification, no human-oversight decision, no source-legitimacy attestation. A RELIABLE narrator grade is an operator-assigned judgment, not a verified fact about the upstream source. See the evidence architecture offer →

The Living Chain#

DESTRUCTIVE ▼ GENERATIVE ▲ GENERATIVE ▲ source RELIABLE pdf- scraper RELIABLE ingest- model UNGRADED answer- model RELIABLE claim CAP Grade ṣaḥīḥ Action SERVE WITH CAVEAT

Every claim carries its chain.

source
RELIABLE
pdf-scraper
RELIABLE
DESTRUCTIVE ▼
ingest-model
UNGRADED
GENERATIVE ▲
answer-model
RELIABLE
GENERATIVE ▲
ḥasan
SERVE WITH CAVEAT
One ungraded narrator caps the grade.
Interactive

Try the Decision Matrix#

Select a chain grade and content verdict — the matrix routes your claim.

Route
SERVE WITH CAVEAT
Show full 4×3 decision matrix
ConsistentContradictionUnverifiable
Ṣaḥīḥ-tierServe directly; cacheʿIlal flag → human reviewServe with caveat
Ḥasan-tierServe with caveat; seek corroborationHold; do not serveHold in review queue
Ḍaʿīf-tierHold; seek corroborationQuarantineHold in review queue
Mawḍūʿ-tierReject & quarantine narratorReject & quarantine narratorReject & quarantine narrator
Interactive

Corroboration & the Madār#

Two chains, same claim. Does corroboration fire — or does the madār block it?

Chain 1 — Base (Ḍaʿīf) src-A scrape-A ingest-A ḍaʿīf-tier ✓ Narrator identity ✓ Model family ✓ Upstream source Chain 2 — Corroborating (Ṣaḥīḥ) src-B scrape-B ingest-B ṣaḥīḥ-tier Independence: 3/3 ✓ CORROBORATION FIRES ḍaʿīf → ḥasan (capped)
Honest annotation: Validated 603/603 on Wikipedia, 104/104 on physics textbooks, 8/8 negative controls. But no negative controls were run on the physics corpus, and both Wikipedia corpora are Wikimedia projects — so v2 validates the mechanics of corroboration, not the independence assumption. Content-level correlated error is not addressed at all.
Section 1

The Problem#

When a multi-agent AI system answers you, that claim has passed through many hands — a source, a scraper, an ingestion model, a synthesis model. Each hand can drop, distort, or invent.

A single factual claim may pass through document retrieval, extraction, summarization, knowledge integration, and answer synthesis before reaching a user. Provenance systems can reconstruct that chain after the fact. They provide limited guidance on the question users actually care about: this specific claim, having passed through these specific transformers — should I trust it?

Provenance tools record what happened. Nothing grades who did it — or tells you how much to trust the result. A confident model at the end of a chain cannot repair a corrupted extraction at the start of it.

Section 2

The Idea#

Classical Islamic hadith science spent twelve centuries on a structurally identical problem: whether to trust knowledge transmitted through chains of human narrators. ISNAD transfers that methodology to AI pipelines.

1

Every claim carries its transmission chain (isnād)

2

Every transmitter keeps a living, per-domain grade (rijāl)

3

The weakest link caps the chain's trust

4

Independent corroboration upgrades — capped and gated (mutābaʿāt)

5

Content is criticized separately from the chain (matn)

Section 3

Concept Mapping: From Ḥadīth Science to AI Pipelines#

The intellectual heart of the framework — the paper's Table 1, mapping each classical term to its ISNAD usage.

TermClassical meaningUsage in ISNAD
isnādThe complete chain of transmitters attached to a narrationThe agent chain attached to a claim
matnThe transmitted content itselfThe claim text
rijāl"The transmitters"; the discipline of evaluating themThe narrator registry
ʿadālahA narrator's integrity / uprightnessResistance to manipulation; source trustworthiness
ḍabṭA narrator's precision and retention accuracyMeasured error rate of an agent/model
al-jarḥ wa-l-taʿdīlFormal criticism and accreditation of narratorsThe registry update process (evals, audits)
ittiṣāl / munqaṭiʿChain continuity / a broken chainComplete trace / missing trace segment
ṣaḥīḥ, ḥasan, ḍaʿīf, mawḍūʿSound, good, weak, fabricated (grades)Trust tiers for claims
mutābaʿātIndependent corroborating chainsMultiple disjoint agent chains asserting one claim
ʿilalHidden defects: sound-looking chain, defective contentChain passes but content criticism fails
muḥaddithThe expert scholar who adjudicates difficult casesThe human reviewer in the loop

Refining the Weakest-Link Rule

In hadith transmission a weak link is unconditionally corrupting. AI transformations are not all alike. A destructive transformation — extraction, chunking, lossy summarization — can only lose or corrupt information, and there the strict minimum applies: nothing downstream recovers what the extractor dropped. A generative transformation can either repair upstream noise or introduce fresh corruption by favouring its parametric prior over retrieved evidence. For generative steps the rule is therefore not a blind minimum but a bounded one: a generative narrator can raise the floor only up to its own grade, and only when corroboration supports the repair — and it can always lower it.

This mirrors a distinction the muḥaddithūn drew themselves, between riwāya bi-l-lafẓ (transmission of exact wording) and riwāya bi-l-maʿnā (transmission by meaning). The weakest-link principle survives; it is refined by transformation type.

Domain-conditioned grading. The registry key is (narrator, domain), not narrator alone. A model precise on historical dates may be unreliable on quantum-mechanical equations. This is a return to classical practice — rijāl scholars graded transmitters domain-specifically as a matter of course.
Version-bump resets. A model-version bump resets a narrator to ungraded until re-evaluated. Version drift is treated as a new narrator, not inherited reputation. Known blind spot: a hosted model whose weights change without a version-string change will silently retain a grade it no longer merits. Deployments against third-party APIs should treat scheduled re-evaluation as mandatory.
Section 4

The Decision Matrix#

The framework's most actionable artifact — chain grade × content verdict → action. The third column is the binding constraint on real-world coverage.

Matn: consistentMatn: contradictionMatn: unverifiable
Ṣaḥīḥ-tier chain
complete; all narrators graded reliable
Serve directly; cacheʿIlal flag → human review. Highest-value review case.Serve with caveat
Ḥasan-tier chain
complete; ≥1 narrator ungraded or mid-tier
Serve with explicit confidence caveat; seek corroborationHold in review queue; do not serveHold in review queue
Ḍaʿīf-tier chain
weak narrator, or munqaṭiʿ
Hold; seek corroborating chain before servingQuarantineHold in review queue
Mawḍūʿ-tier chain
rejected narrator — e.g. known injection source
Reject and log; quarantine the narratorReject and log; quarantine the narratorReject and log; quarantine the narrator
The third column is not a formality. Content criticism is a partial function. A critic that cannot evaluate a claim must be able to say so rather than defaulting to "consistent." Unverifiable is the classical position of tawaqquf — suspension of judgment. Routing it conservatively is what makes the framework safe under a weak critic. It is also the binding constraint on practical coverage. A framework that omitted this column would look far more capable on paper and behave far worse in deployment.
Sound chain + contradiction is the most informative signal in the system, not an error state. It means either a trusted source has changed the world's state, or the corpus contains a latent defect. Both are exactly what a self-maintaining knowledge base exists to catch.
The mawḍūʿ tier is the native defence against knowledge-base poisoning. A compromised source or manipulated agent is, in rijāl terms, a fabricator — a narrator whose ʿadālah has failed. Grading for integrity and not merely precision means the framework quarantines an injection source rather than merely noting low confidence in its output.

Contradictions go to humans by default. Knowledge-conflict research consistently finds LLMs unreliable at representing and reconciling competing evidence, and during prototype development LLM synthesis handled contradiction cases worse than an extractive baseline until contradiction markers were made explicitly visible in context — after which auto-resolution was still left disabled in favour of gated review.

Section 5

Worked Example: End to End#

Claim: "the momentum of a photon is p = h/λ", ingested from an openly licensed physics text.

Chain: [OpenStax University Physics Vol. 3 — source; publisher-trusted, ʿadālah high] → [pdf-scraper v1.2 — ḍabṭ graded: high extraction fidelity on text-layer PDFs] → [ingest-analysis model M@v — ungraded, recent version bump] → [ingest-renderer M@v — ungraded]

Chain grade: complete (ittiṣāl holds), but two narrators ungraded → ḥasan-tier by the weakest-link rule.

Matn criticism: the corpus already contains a page asserting momentum as p = mv without regime qualification. Contradiction detected.

Matrix cell: ḥasan × contradiction → review queue; do not serve.

Adjudication: a human reviewer recognises both claims as valid in different regimes — classical and quantum — and updates both pages with regime qualifiers, resolving the contradiction once, at ingest time, rather than leaving every future query to renegotiate it. The resolution feeds back: the ingestion model's contradiction-detection behaviour on this case becomes taʿdīl evidence toward its grade.

This example is implemented as a runnable demo and an integration test in the reference implementation (make demo, examples/worked_example.py, tests/test_worked_example.py).

Section 6

What ISNAD Adds Over Existing Provenance#

Each Mechanism Against Its Strongest Existing Ancestor

MechanismStrongest ancestorWhat ISNAD adds
Chain aggregationJøsang subjective logic (2001) — trust discounting along chainsTransformation typing (riwāya bi-l-lafẓ vs. bi-l-maʿnā); completeness semantics
Source registryTruth discovery (Yin et al. 2008; Dong et al. 2014)Per-domain grading; version-bump resets; quarantine lifecycle
Content criticismFEVER-style fact-checkingDecoupled from chain quality; operates on matn independently; explicit unverifiable verdict
Agent gradingMAS reputation (EigenTrust, ReGreT, FIRE)Claim-level provenance, not just partner selection
Decision routing— (no direct ancestor)Chain grade × content verdict → concrete serve/review/quarantine action

Capability Coverage

ApproachChain captureGraded transmittersContent criticismDecision routing
W3C PROV / PROV-AGENT✅❌❌❌
Truth discovery (TruthFinder, KBT)❌✅partial❌
MAS reputation (EigenTrust, ReGreT)❌✅❌partial
FEVER-style fact-checking❌❌✅❌
ISNAD✅✅✅✅

Chain capture means claim-level transmission chains are recorded per claim; graded transmitters means each transformer carries a maintained reliability grade; content criticism means content is evaluated independently of chain quality; decision routing means the system emits a concrete serve/review/quarantine action. A ✅ indicates the approach addresses that capability as a first-class concern.

Deep lineage: Condorcet's jury theorem (1785), Hume on testimony (1748), the hearsay-within-hearsay rule (US FRE 805), Dempster–Shafer evidence combination, Dawid & Skene (1979), Jøsang's subjective logic (2001), TruthFinder (2008), Knowledge Vault (2014), Knowledge-Based Trust (2015). The paper's claim is not novelty for its own sake but the transfer of the most operationally complete solution to a domain that lacks it.

Section 7

Quickstart#

Current release: v2.24.0 (released 11 Sep 2026) · Python 3.11+ · Apache-2.0 · 985 tests
Hosted ISNAD Cloud — the same core, run for you at isnad.islamandai.com: self-host free (Apache-2.0) · Pro $79/mo (50k claims/mo, 90-day retention, PDF audit reports) · Business $499/mo (500k claims/mo, SSO/SAML/SCIM + RBAC) · Enterprise from $5k/mo · one-time Trust Assessment $5,000. The open-source core stays Apache-2.0 forever.

Install

pip install isnad
npm i isnad

Also on npm: the JS/TS audit-record verifier — verifies SHA-256 / Merkle / detached-signature integrity of ISNAD audit records. Grading stays in the Python core.

60-Second Example

from isnad import Registry, Chain, ChainLinkSpec, grade_chain, decide
from isnad.types import NarratorGrade, ContentVerdict
from isnad.critics import EmbeddingCritic

# Build a chain: source → scraper → model
chain = Chain([
    ChainLinkSpec("openstax-textbook", 0, domain="physics"),
    ChainLinkSpec("pdf-scraper-v2", 1),
    ChainLinkSpec("ingest-model-v3", 2),
])

# Seed-grade known narrators
reg = Registry()
reg.register("openstax-textbook", "physics", grade=NarratorGrade.RELIABLE)
reg.register("pdf-scraper-v2",    "physics", grade=NarratorGrade.RELIABLE)
reg.register("ingest-model-v3",   "physics", grade=NarratorGrade.ACCEPTABLE)

# Grade the chain
grades     = [reg.get_grade(l.narrator_id, l.domain) for l in chain.links]
transforms = [l.transform_type for l in chain.links]
chain_grade = grade_chain(grades, transforms, is_complete=True)

# Content criticism — decoupled from chain grading
critic  = EmbeddingCritic()
verdict = critic.evaluate("p = h/λ", "p = h/lambda", ["p = mv"])
action  = decide(chain_grade, verdict)

print(f"Chain: {chain_grade.value.upper()} | Content: {verdict.value} | Action: {action.value}")

LangChain Integration (5 Lines)

pip install isnad[langchain]

from isnad.integrations.langchain import IsnadTracer, seed_registry
from isnad.critics import EmbeddingCritic

reg = seed_registry({"source:docs": "reliable", "model:gpt-4o": "acceptable"})
tracer = IsnadTracer(registry=reg, critic=EmbeddingCritic())
chain.invoke("What is F=ma?", config={"callbacks": [tracer]})
print(tracer.report())

MCP Integration — Grade MCP Servers as Narrators

from isnad.integrations.mcp import MCPToolObserver, build_mcp_tools, handle_grade_claim

obs = MCPToolObserver(registry, domain="finance")
obs.record_call(tool_name="query", server_name="data-warehouse")
result = obs.grade("revenue grew 12% YoY")
print(result["chain_grade"])   # weakest-link grade over the tool narrators

The Model Context Protocol treats every tool as a transmitter of claims. ISNAD grades MCP servers as narrators — tool outputs are destructive links, so tools stay majhūl (UNGRADED → capped at ḍaʿīf) until the operator seeds or logs evidence. A grade_claim tool exposes the operator's registry to agents at runtime — grades only, never manufactured.

Run the Experiments

git clone https://github.com/alizahidraja/isnad && cd isnad
make install   # uv sync
make test      # zero config, SQLite fallback
make demo      # the paper's worked example (§4.5)
make check     # lint + type-check + test

No database required for pure-logic tests; PostgreSQL is optional (docker compose up, set ISNAD_DATABASE_URL).

Section 8

Architecture & Module Map#

ConceptWhat it doesModule
isnād (chain)Ordered, gap-checked transmission chain per claimisnad/core/chain.py
rijāl (registry)Graded narrator store per (alias@version, domain)isnad/core/registry.py
jarḥ–taʿdīlEvidence-driven state machine for narrator gradesisnad/core/registry.py
Bayesian gradingBeta-distribution narrator grades (default)isnad/core/registry.py
ittiṣālCompleteness as epistemic property (gap → ḍaʿīf)isnad/core/chain.py
Weakest-link gradingChain grade = refined minimum over narratorsisnad/core/grading.py
mutābaʿāt (corroboration)Independent-chain upgrade + madār detectionisnad/core/corroboration.py
matn criticismContent evaluated independently of chain qualityisnad/critics/
Decision matrixChain × content → action routerisnad/core/decision.py
PersistenceSQLAlchemy-backed registry (swap via protocol)isnad/storage/
APIFastAPI service with DI + Prometheus metricsisnad/api/
CLIisnad serve, isnad seedisnad/cli/
ʿadālah / ḍabṭIntegrity and precision as two distinct axesisnad/types.py
ISNAD system architecture — layered diagram of types, core engine, content critics, and audit/API/integrations
System architecture (layered view). The full 3-tab diagram — including the Validation Matrix with the κ results — lives at docs/ARCHITECTURE.drawio →

Pluggable Strategies — Updated Defaults

StrategyProtocolDefaultWhat it decides
GradingStrategyisnad/types.pyRefinedWeakestLinkHow link grades combine into a chain grade
TransitionPolicyisnad/types.pyBayesianTransitionPolicyHow evidence moves narrators between states
CorroborationPolicyisnad/types.pyCappedCorroborationPolicyHow independent chains upgrade a claim
CorrelationDetectorisnad/types.pySharedLineageDetectorWhether two chains are truly independent
ContentCriticisnad/critics/base.pyHybridCritic / EmbeddingCriticContent contradiction detection

Reference Schema

CREATE TABLE rijal_claims (
    claim_id         TEXT PRIMARY KEY,
    page_slug        TEXT NOT NULL,
    claim_text       TEXT NOT NULL,
    narrator_chain   JSONB NOT NULL,
    chain_confidence NUMERIC(4,3),
    valid_from       TIMESTAMPTZ,
    valid_until      TIMESTAMPTZ,
    superseded_by    TEXT,
    chain_status     TEXT DEFAULT 'active'
);

CREATE TABLE narrator_registry (
    narrator_id      TEXT NOT NULL,
    domain_tag       TEXT NOT NULL,
    narrator_type    TEXT NOT NULL,
    grade            TEXT NOT NULL,
    known_error_rate NUMERIC(4,3),
    model_version    TEXT,
    is_active        BOOLEAN DEFAULT TRUE,
    PRIMARY KEY (narrator_id, domain_tag)
);

Because the pipeline's transformation points already emit trace identifiers, populating narrator_chain is an instrumentation task, not an architectural change.

Section 9

Evaluation: What's Validated, What Failed, What's Measured#

ISNAD-Bench — Measured Against 1,200 Years of Ground Truth

The framework's newest and strongest validation. ISNAD-Bench grades 577,024 real hadith chains that classical scholars already graded by hand, then asks one falsifiable question: when ISNAD is given the scholars' own narrator grades, does its weakest-link rule reach the same chain verdicts they did?

The answer is reported as Cohen's κ — never raw accuracy, which is meaningless on a corpus dominated by ṣaḥīḥ.

ComparisonCohen's κAgreement
ISNAD vs scholarly consensus (strict default)0.871489.7%
ISNAD vs consensus (lenient opt-in)0.761082.4%
Human ceiling — a single scholar vs consensus0.450—
Human ceiling — scholar vs scholar0.331—
Shuffled-rank control0.047—
Corpus577,024 scholar-graded chains
The number that matters is 0.331, not 0.871. The scholars agree with each other at κ = 0.33 — the ground truth itself is contested. Hadith criticism is an argument, not a lookup table. What 0.871 says is smaller and more useful: ISNAD faithfully implements the scholars' method — it is not better than the scholars. Same inputs, same verdict the tradition converges on, reproducibly, at scale, no human in the loop.
Negative controls. Shuffle the narrator grades and κ collapses to 0.047; a majority-class predictor scores 0.000. The signal is in the grades, not an artifact of the pipeline.

Three things the benchmark taught

  1. Mutābaʿa was already load-bearing. The largest "disagreement" was chains scholars called weak-alone-but-ḥasan-if-corroborated. ISNAD said weak; they said acceptable. Checking the corpus, 88–92% of those chains have an independent corroborating route — the scholars were applying mutābaʿa, and it validated ISNAD's corroboration engine from the outside.
  2. Reliability has a timeline. 161 narrators declined over their lifetime (ikhtilāṭ), touching 35% of all chains. The scholars didn't discard them — they dated the decline. ISNAD's get_grade_as_of() period-slicing, built on principle, turns out to be load-bearing across a third of the corpus.
  3. The default was wrong. When a narrator had no grade (majhūl), ISNAD treated them as passable — generous, reasonable, wrong. κ = 0.761. The classical position caps unknown at ḍaʿīf. Flipping the default: κ = 0.871. Eleven points of κ from a 1,200-year-old method being right where modern intuition was wrong.

Full results, preregistered mapping, negative controls and the complete disagreement buckets: bench/docs/RESULTS.md · preregistered mapping.

RAGTruth — Hallucination Detection on Real LLM Output

A second, independent benchmark. On RAGTruth — 17,790 responses across 6 models — ISNAD's grounding critic catches 97.6% of hallucinations at κ = 0.575 (79.3% accuracy against a 71.4% baseline). Reported on ISNAD Cloud with the same preregistered discipline as ISNAD-Bench.

Setup

  • 20,000 atomic claims extracted from real PDF text: OpenStax University Physics Vols 1–3 (CC BY 4.0) and Crowell's Light and Matter (CC BY-SA), split evenly, 10,000 per source.
  • Domain tags: mechanics (7,600), general (6,890), electromagnetism (3,054), modern-quantum (1,728), optics-waves (728).
  • Extraction via deepseek-chat over chunked PDF text.
  • Split 30% calibration / 70% evaluation → 14,001 evaluation claims, across 10 random seeds.
  • Four narrators with designed fault rates: pdf-scraper@1.2 (1%), ingest@good (2%), ingest@weak (15%), pdf-scraper@0.9-legacy (18%).
  • Faults: OCR noise, digit swap, sign flip, negation drop, unit corruption, formula mangling, entity swap, fabricated numerics, regime confusion. ~5% of chains marked incomplete.
  • Review budgets: B ∈ {2%, 5%, 10%, 20%}.
Definitions. Error: a served claim counts as an error if its text was altered by fault injection, per the manifest — synthetic corruptions of correct source text, not naturally occurring factual errors. Coverage: the fraction of evaluation claims the decision matrix routes to serve or serve-with-caveat. Claims routed to review beyond budget, and quarantined claims, do not count.

Status at a Glance

ComponentStatusWhat remains
Weakest-link grading and quarantine✅ Validated—
jarḥ–taʿdīl grade recovery⚠️ PartialRecovered 3/4; missed the highest-fault narrator through insufficient per-cell calibration evidence
Corroboration (mutābaʿāt)✅ ValidatedNegative controls on the physics corpus; organically weak chains
Confidence baseline⛔ Not testedBaseline is synthetic; real model self-confidence untested
Matched-coverage (§8.4)✅ MeasuredChain-only 47.9% coverage @ 3.0% served-error; critic adds nothing on subtle corruption
Content criticism✅ MeasuredLLM 1.000 recall / 0.000 false-consistent; offline tiers conservative
End-to-end serving pipeline⚠️ Partiale2e utility measured (0.0% vs 50.0% served-error); live deployment open

Validated: Weakest-Link Quarantine

Every claim whose chain contained a narrator graded rejected was quarantined — 4,057 claims, 29% of the evaluation split — and every quarantine was traceable to the specific narrator grade that caused it.

Step 0: source:openstax RELIABLE Step 1: pdf-scraper@1.2 RELIABLE [DESTRUCTIVE] Step 2: ingest@weak REJECTED [GENERATIVE] <- binding link Chain grade: MAWDU -> REJECT_AND_QUARANTINE_NARRATOR

Partial Failure: The jarḥ–taʿdīl Loop

The loop recovered three of four narrator grades from audit evidence alone:

  • pdf-scraper@1.2 (1% fault) → reliable in all 50 (narrator, domain) cells
  • ingest@good (2%) → acceptable in all 50
  • ingest@weak (15%) — deliberately left ungraded — independently discovered and driven to rejected in 49/50 cells, weak in one
The fourth was not recovered, and the failure is the most instructive result in the paper. pdf-scraper@0.9-legacy, carrying the highest designed fault rate (18%), never accumulated enough claims in the calibration split to earn any grade in any domain, and remained ungraded throughout. It was among the largest single sources of injected faults. The loop found the unreliable narrator it was given enough evidence about, and silently missed the one it was not.

Grade recovery is bounded by calibration coverage per (narrator, domain) cell, not only by the transition policy. Domain-conditioned grading multiplies the number of cells that must be filled, and a narrator rare within a domain is invisible to jarḥ–taʿdīl regardless of how unreliable it is. A deployment should monitor per-cell evidence counts as a first-class operational metric and treat a persistently under-evidenced cell as a risk, not a neutral absence.

Validated: Corroboration Across Three Corpora

v1 (exact)v2 (Wikipedia)v3 (physics)
Matching criterionexact stringcosine ≥ 0.75cosine ≥ 0.80
Corpus12 article intros30 article pairs2 textbooks
Sentences21510,54422,372
Candidate pairs136662104
Pairs evaluated136603104
Match density—6.3%0.5%
Corroboration fired68/136603/603104/104
Negative controlsnone run8/8none run
Independence basissynthetic chainsseparate editorial communitiesdifferent authors and publishers
Difficultyeasymediumhard
Important caveats. v2 independence is imperfect — English and Simple English Wikipedia have separate editorial communities but share the Wikimedia umbrella. v2 validates the mechanics of corroboration rather than the independence assumption. v3 is the stronger independence test (different authors, publishers, decades, and upstream domains). No negative controls were run on v3 — a 100% fire rate without a discrimination check is weaker evidence than the same rate with one. In both v2 and v3 the weak-tier baseline chain was created by assigning one narrator a weak grade because the sources are high quality. The experiment tests whether corroboration upgrades a weak chain when an independent chain corroborates it — not whether it works on chains that are weak for organic reasons.

The eight negative controls (v2) correctly produced no upgrade: no matching text; shared model family (the madār case); corroborators below the grade gate; a rejected base chain; the ṣaḥīḥ cap; shared upstream source; insufficient independent chains; empty corpus.

Matched-Coverage: The Measured Finding

The paper reported the matched-coverage comparison as "inconclusive" — ISNAD-gated coverage could not be driven above 4.8% with the reference critic. That question is now resolved as a measured, three-part finding (#126), not a shrug.

1. Chain grading + jarḥ–taʿdīl is the primary gate, and it works. With the critic treated as a no-op (acceptance riding entirely on the chain and narrator grades), ISNAD reaches 47.9% coverage at 3.0% served-error — the corrupt narrator is caught by the evidence loop, not by a content critic.

2. The matn critic is a rejection gate, not an acceptance gate. On a sample of the residual "leaks" (corrupted claims with a sound chain), the LLM critic caught 0 of 30 — subtle single-claim corruption is exactly what a content critic cannot reliably detect. Its false-consistent rate on meaning-changing corruption is 39.1% (#126), so it is unsafe as an acceptance gate.

3. The missing discriminator — corroboration — is structurally absent from this corpus. The classical defense against "sound chain, subtly wrong claim" is independent corroboration (mutābaʿāt). The §8 corpus has zero cross-source overlap, so corroboration fired 0 times by construction. Matched coverage on this corpus is structurally unattainable, not merely unfinished.

Full write-up: experiments/s8_gated_vs_ungated/results/MATCHED_COVERAGE_FINDING.md.

End-to-End Utility: The Fidelity→Utility Crossing

Every prior validation used synthetic fault injection or extracted sentences. Here the claims are generated by a live LLM (DeepSeek), labelled by a deterministic oracle (not the critic), then routed through the real pipeline.

served-error
ISNAD-gated0.0%
Ungated50.0%

The gate served only the correct in-corpus restatements (with caveat) and routed all 12 novel-topic claims to review.

Frame it honestly: this is a mechanism demonstration — one corpus, one model, one critic, one run — not a hallucination-detection benchmark. The novel topics are genuinely out-of-corpus, so "review" is the easy, correct call. The hard case — a confident wrong claim within a covered domain — is not what this experiment stresses.

Full write-up: experiments/e2e_utility/RESULTS.md.

The Confidence Baseline Is Uninformative by Construction

A confidence-gated condition was included to represent current practice. It performs at chance — but for a reason that sharply limits what can be concluded from it. The corpus's model_confidence field is computed before fault injection and never updated. Confidence is statistically independent of corruption by construction. The point-biserial correlation is r = 0.04; mean confidence on corrupted claims (0.853) and clean claims (0.848) differ by less than one hundredth.

This is a floor, not a finding. It establishes that the ISNAD condition is not being measured against a baseline that trivially wins, and nothing more. No claim is made about whether real LLM self-confidence carries a defect signal.

Honest Negatives: Transition Policy & Content Critic

Sweeping the downgrade threshold shows a direct tradeoff with no sweet spot:

Downgrade thresholdCoverageServed errorReview precisionGrades recovered
37.1%0.0%0.112/4
69.4%0.0%0.152/4
1010.0%0.0%0.111/4
1510.0%0.0%0.130/4
2510.0%0.0%0.130/4

There is no setting in this range that delivers both coverage and grade recovery. More audit evidence did not help — increasing the per-cell audit budget reduced coverage from 10.0% to 5.9% at 40 and above. The default policy over-penalises reliable narrators as evidence accumulates on small samples.

The content critic is the binding constraint — and now it's measured. The critic tiers are documented with their measured behaviour, not marketed as equal:
CriticContra. recallFalse-consistent
LLMCritic (DeepSeek)1.0000.000
LocalNLICritic0.7600.000
HybridCritic0.7200.040
EmbeddingCritic0.1200.320

Only the LLM tier is near-perfect on genuine contradictions (100% recall, zero false-consistents); the offline tiers are safe but conservative. But on subtle meaning-changing corruption, the critic's false-consistent rate is 39.1% (#126) — it is a rejection gate, not an acceptance gate.

The earlier "~90% LLM / ~40% NLI / ~30% LocalNLI" figures were estimated on a template-injected word-swap set and have been retracted (issue #96); the measured hierarchy above replaces them.

The honest negative stands: with most real claims returning "unverifiable," the critic — not the grading mechanism — caps end-to-end coverage. The hierarchy is measured; the open piece is the hard case (a confident wrong claim within a covered domain).

Case Study: Contradiction Detection on Real Texts

A prototype compiled-wiki system ingested OpenStax University Physics Vols 1–3 and Crowell's Light and Matter via chunked PDF extraction, with DeepSeek-chat as the ingestion model, single run.

Result: 19 contradictions surfaced across the corpus, each manually reviewed and confirmed genuine — among them classical momentum (p = mv) versus photon momentum (p = h/λ); wave versus particle treatments of light; Newtonian versus relativistic kinematics.

All 19 were Type B — regime distinctions where both claims are true in their domains and resolution means qualification — not Type A errors. This validates the detection substrate and routing path; it does not demonstrate error-correction capability.

Four important caveats: (1) Not reproducible — the prototype is a proprietary system, not part of the open-source release. (2) Precision, not recall — 19/19 confirmed genuine, zero false positives. "19 contradictions found" means "at least 19 exist and were caught," not "exactly 19 exist." (3) All Type B — regime distinctions, not errors. (4) Parametric knowledge confound — the ingestion model already knew undergraduate physics. Every contradiction was confirmed to correspond to explicit statements in the ingested chunks, but the experiment does not isolate retrieved evidence from parametric knowledge.
Section 10

Limitations#

The analogy has edges.

Classical narrators were moral agents whose ʿadālah was a judgment of character formed over a lifetime; models are stochastic artifacts whose "integrity" is an engineering property of their deployment. The mapping is functional, not metaphysical. Classical grading also rested on scholarly consensus formed over generations; a registry populated by one team's evaluations is a far thinner epistemic base, and its grades should be presented with corresponding humility.

Registry cold start.

Weakest-link grading with an empty registry caps everything at ḥasan-tier or below. Arguably correct — new pipelines should earn trust — but it front-loads human review cost. Practical value depends on the jarḥ–taʿdīl loop converging faster than review budgets exhaust. §8 showed the intended mitigation (seed grading) goes less far than anticipated.

Grading humans.

Where human contributors are narrators, the registry becomes a reputation system over people, with the fairness, contestability, and workplace-power questions that implies. Flagged as requiring governance design — contestable grades, transparent criteria — not resolved.

False precision.

chain_confidence as a decimal invites over-trust. The ordinal-first principle is the mitigation; implementations that surface raw numbers to end users without tier context would be misusing the framework.

Independence is an idealisation.

Two chains may share no explicit narrator yet still fail together — if both terminate in the same answer model, or in models from a family with correlated training data. Classical scholars faced the same problem in the madār (the pivot narrator through which ostensibly independent chains converge). The reference implementation enforces an independence check on narrator identity, model family, and upstream source — but a genuinely correlated pair sharing none of those three would still pass, and content-level correlated error (two independent sources repeating the same received mistake) is not addressed at all.

Probabilistic and provisional claims.

The framework assumes claims with determinate truth values. "The Higgs boson mass is 125.1 GeV" carries an implicit uncertainty interval. Contradiction detection between a point estimate and an interval is not defined. The lifecycle columns provide a mechanism for supersession but not a semantics for uncertainty.

Storage and query overhead is estimated, not measured.

Chains are short (typically 3–5 links) so per-claim JSONB is small and bounded; the registry grows with pipeline changes, not corpus size. At the targeted scale — thousands to low millions of claims — neither table is expected to bottleneck. These are design expectations, not measured results. No latency or storage benchmark was run, and none is claimed.

Scope exclusions: Byzantine collusion among agents, cryptographic attestation of traces (complementary, not competing), and open-ended or creative generation. ISNAD grades claims that can in principle be true or false and checked against a corpus. It is a framework for knowledge accumulation, not for tasks where there is no fact of the matter to transmit.

Positionality. I am a Muslim practitioner and builder. ISNAD is where two lines of my work merged: Islam & AI, the Islamic knowledge platform I founded, and my work engineering multi-agent AI systems. This work adapts a methodology developed by hadith scholars; it makes no religious claims, issues no rulings, and does not position any computational system as an authority within the religious sciences themselves. The direction of intellectual credit should be unambiguous: the framework's rigor belongs to twelve centuries of muḥaddithūn; the transfer to AI systems is the contribution claimed here. I would welcome collaboration with hadith-studies scholars to correct any misrepresentation of the classical discipline.
Section 11

What's Next#

Shipped Since the Paper's Submission Version

  • ISNAD-Bench — measured against 577,024 real hadith chains. Cohen's κ = 0.8714 vs scholarly consensus (strict, 89.7% agreement), 0.7610 (lenient), shuffled control 0.047, human ceiling 0.331. Preregistered mapping + negative controls.
  • Strict default — ungraded narrators (majhūl) now cap at ḍaʿīf, matching the classical position; leniency is opt-in via lenient_unknown=True.
  • Multi-provider LLM critic — OpenRouter, OpenAI, DeepSeek, Anthropic, Gemini, Groq, Together, or local Ollama; measured 1.000 recall / 0.000 false-consistent on the committed eval set.
  • Audit evidence layer — tamper-evident AuditRecord with RFC 8785-canonical SHA-256, DAG support, linear hash chain + Merkle batch log, detached signatures (HMAC + Ed25519), PII redaction, and governance mapping.
  • Period-sliced grades (get_grade_as_of()) — the ikhtilāṭ remedy, validated by ISNAD-Bench M4.
  • Per-role precision (#3) — precision (ḍabṭ) graded per (narrator, role, domain); integrity stays per-narrator. Integrity ladder: ʿadālah never forgets, ḍabṭ recovers.
  • Evidence-backed seed grades — Registry.seed() / seed_from_benchmark() record evidence, not bare grades; seeds survive the Bayesian posterior.
  • Chain independence (#125) — document-hash madār detection and a disjointness discount; corroboration is weighted by independence score, and provenance is surfaced in results/API.
  • Working content critics: EmbeddingCritic, HybridCritic (NLI), LLMCritic
  • LangChain integration: IsnadTracer + seed_registry + IsnadMiddleware (gates claims); plus LangGraph, CrewAI, LlamaIndex, OpenTelemetry ingest.
  • MCP integration — grade MCP servers as narrators: MCPToolObserver records tool calls as TOOL links, and a grade_claim tool exposes the operator's registry to agents (duck-typed; no MCP SDK required).
  • JS/TS audit-record verifier on npm — npm i isnad verifies SHA-256 / Merkle / detached-signature integrity of audit records (verifier only; grading stays in Python).
  • FastAPI service, CLI, SQLAlchemy persistence, endpoint/version-drift handling
  • Live Verify integration — consume a cryptographic seal as a high-trust narrator
  • Model-drift leaderboard (#71) — a preregistered, LLM-free harness measuring hallucination across chain depth 1–5, with a deterministic offline drift injector and negative controls. First live result: hallucination originates at memory-generation (hop 1) and stays flat across depth — the relay preserves the error, it doesn't compound it.
  • Chain-scoped content grounding (#216) — chain_scoped_corpus() + a swappable GroundingPolicy flag a claim grounded only off its own chain (a grounding gap, not proven contamination).
  • Content-madār tightening (#232) — shared-error fingerprinting now requires error-distinctive signals (number+unit, citation+number/date); false-positive-on-agreement down from 0.750 to 0.375 with token-bearing recall held at 1.0.
  • Warm default registry + isnad scan (#203/#204/#206) — an evidence-sourced seed set (~39 narrators) so a fresh pipeline no longer starts fully ungraded, plus a CLI that maps live transmitters onto the registry (vouched vs cold).

Highest-Priority Open Work (collaborators welcome)

  1. The hard case. ISNAD-Bench validates the deterministic grading rule, and the end-to-end utility experiment (#128) shows the gated pipeline serving 0.0% vs 50.0% served-error on live LLM-generated claims. What remains is the hard case: a confident wrong claim within a covered domain, where "review" is not the easy call.
  2. A live production deployment where chains are written automatically at ingest and the registry gates a serving path under real query load.
  3. Chain independence — the hardest open problem: topology cannot prove two chains are independent. Tracking issue →
  4. Negative controls on the physics corroboration corpus (v3) and corroboration tested on organically weak chains.
  5. Per-cell evidence-count monitoring as a first-class operational metric.
  6. Extension to probabilistic and provisional claims.

Good First Issues

  • An alternative critic (sentence-transformers embedding, CrewAI integration)
  • A seed-grade bootstrapper from published benchmark accuracies
  • A new CorrelationDetector using embedding similarity or model-card lineage
  • Extend semantic corroboration to multi-source corpora
  • A calibrated TransitionPolicy from your own pipeline's §8 experiment data
  • Domain-specific ContentCritic with formula canonicalisation
  • Pipeline adapters for CrewAI or AutoGen tracing
Section 12

In the Wild#

Live Verify — Independent Comparison

Paul Hammant — co-creator of Selenium (~22 years ago), maintainer of trunkbaseddevelopment.com, ex-ThoughtWorks, based in Scotland. Emailed Ali directly after finding the repo: "i saw your isnad repo, and looks a very interesting approach.. i'd love contributing to it incase i have the time."

A dedicated comparison document: Live Verify vs. Isnād–Rijāl

Both attack the same root problem — why should I trust this claim that reached me through a chain of hands? — and both reach for a provenance answer rather than a plausibility answer. But they sit at opposite ends of the trust pipeline. They are complementary rather than competing.

ISNADLive Verify
LayerInside an AI pipeline, per-claimBetween organisations, per-artifact
MechanismGraded reputation registry + Bayesian loopCryptographic hash + domain endpoint
Trust rootAccumulated narrator track recordAn authorizedBy chain to a jurisdiction's root namespace
Outputserve / review / quarantineverified / not-verified + status
Truth stanceProbabilistic (how reliable)Binary (unaltered vs. altered)
HandlesDistortion by AI transformersForgery/tampering of documents
Both frameworks are chain-based, but the chains carry different things. ISNAD's isnād is a transmission chain — the ordered list of hands a claim passed through, each graded on how reliably it forwards claims. Live Verify's authority chain is a vouching chain — not who relayed the artifact, but who vouches the issuer is entitled to issue it. So where ISNAD asks "how trustworthy is this path of transmitters?", Live Verify asks "is this issuer a legitimate authority?" Both refuse to stop at the leaf; they just walk in different directions — ISNAD backwards through transmission, Live Verify upwards through authorization.
The composition argument. A Live Verify seal is an ideal high-trust narrator input to an ISNAD chain. A cryptographically-anchored, authority-chained source is exactly the kind of link that deserves a top narrator grade — its integrity axis (ʿadālah) is anchored by cryptography and by an authority chain to a sovereign root, rather than by accumulated track record. A Live Verify seal can bootstrap a narrator to a high grade on day one, before any evidence history exists.
Verify this page's own verdict with your phone. ISNAD now issues its own sealed verdicts as a Live Verify issuer. Select the verdict text at alizahidraja.com/verify/isnad-demo and run a Live Verify client — it resolves to verified / self-verified (amber), not green. A page about provenance that carries its own verifiable provenance is the demo. See the issuer surface →

"ISNAD's README carries a 'What's Validated vs. What's Not' box and reports two of its own results as inconclusive rather than positive. … Both projects treat scoping-out over-claims as a first-class feature, not an afterthought." — from the Live Verify comparison document.

Read the full comparison → · Have your own project or comparison? Add it — reach out.

Listed in LangChain's Middleware Directory

ISNAD ships as a community middleware integration in LangChain's official docs. In the "All middleware" directory — sorted by monthly downloads — it ranks sixth, right after OpenAI, Anthropic, AWS, Microsoft Foundry, and CopilotKit.

Screenshot of LangChain's middleware directory showing ISNAD listed sixth, after OpenAI, Anthropic, AWS, Microsoft Foundry, and CopilotKit
The "All middleware" table, top six entries — LangChain middleware directory. Screenshot Sept 2026.

Indexing, Archiving & Aggregation

Also discoverable via NASA ADS, Google Scholar, Semantic Scholar, Connected Papers, alphaXiv, CatalyzeX, Litmaps, scite, OpenAIRE, and Hugging Face Daily Papers. Reddit launch threads — URLs pending; will be linked when recovered.

Prior & Adjacent Work

  • PROV-AGENT (Souza et al., IEEE e-Science 2025) — extends W3C PROV to capture agent interactions.
  • Agent-Sentry (Sequeira et al., 2026) — bounding LLM agents via execution provenance.
  • HadithRank (Hawramani) — models isnād topology using information theory, treating independent chains as redundant channels. That a mathematician grading hadith and this paper grading agents converge on the same redundancy rule is evidence the mapping is structural, not decorative.
  • The Provenance Paradox (Prakash, 2026) — when delegates can inflate self-reported quality, quality-based routing can perform worse than random. Consistent with ISNAD's design intuition but not a replication.
  • RAPS (2026) — reputation-aware publish–subscribe with Bayesian overlays.
  • TrustTrade (2026) — replaces "uniform trust" with cross-agent consistency weighting.
  • AgentPoison (Chen et al., NeurIPS 2024) — red-teaming LLM agents by poisoning memory/knowledge bases; the demonstrated threat the mawḍūʿ tier defends against.
  • Karpathy's "LLM Wiki" sketch (2026) — the compiled-knowledge-base system class that motivates the whole framework.

Two public repositories independently reference isnad-style provenance chains: cognalith/isnad (agent skill security with provenance verification) and VAMaksimov/threat-intel-format (threat intelligence schema with isnad-style provenance chains). Neither is affiliated.

A note on the name. "ISNAD" here refers to the Isnād–Rijāl Framework described in Grading the Narrators (arXiv:2607.24117). Several unrelated projects in the AI-trust space use the same Arabic word — it is a common noun meaning "chain of transmission" and no one owns it. This project is open-source under Apache-2.0, has no token, no blockchain component, and no commercial affiliation with any similarly named product.
Section 14

Cite#

Paper

@article{raja2026grading,
  author        = {Ali Zahid Raja},
  title         = {Grading the Narrators: An Isnad--Rijal Framework for
                   Claim-Level Provenance in Multi-Agent Knowledge Systems},
  year          = 2026,
  doi           = {10.48550/arXiv.2607.24117},
  eprint        = {2607.24117},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI}
}

Software

@software{raja2026isnad,
  author = {Ali Zahid Raja},
  title  = {Isnad--Rijal Framework: Reference Implementation},
  year   = 2026,
  doi    = {10.5281/zenodo.21216872},
  orcid  = {0009-0003-7875-4590}
}

LLM usage disclosure: drafting and editing were assisted by a large language model; all technical claims, design decisions, experimental results, and final text were reviewed by and are the responsibility of the author.

FAQ

Frequently Asked Questions#

What is ISNAD in one sentence?

ISNAD is an open-source framework that grades every agent, scraper, and model in a claim's transmission chain — adapted from classical Islamic hadith science — to tell you how much to trust a claim that has passed through multiple AI hands.

How is this different from W3C PROV or PROV-AGENT?

PROV records what happened — it captures chains, agents, and activities. ISNAD grades who did it and whether to trust the result. PROV gives you a transmission log; ISNAD tells you which link in that log is the weakest, and what action to take. See the positioning tables for a detailed comparison.

Does ISNAD outperform confidence-based gating?

The matched-coverage question is now a measured finding, not an open one. Chain grading alone reaches 47.9% coverage at 3.0% served-error (no content critic); the matn critic adds nothing on subtle single-claim corruption (0/30 caught); and corroboration — the one mechanism that could have discriminated — is structurally absent from the §8 corpus. See the measured finding.

What does "the chain is graded by its weakest link" mean in practice?

A ṣaḥīḥ model at the end of a chain cannot rescue a ḍaʿīf scraper at the start. The chain's grade is the minimum grade among its narrators — adjusted by transformation type (destructive vs. generative) — because information lost early in the chain cannot be recovered downstream. The chain animation demonstrates this visually.

What is mutābaʿāt / corroboration, and when does it not fire?

Mutābaʿāt is the mechanism by which independent chains asserting the same claim corroborate each other, upgrading the claim's grade (capped at ṣaḥīḥ). It does not fire when chains share a narrator, model family, or upstream source — the classical madār case. It is validated at 603/603 on Wikipedia and 104/104 on physics textbooks, with 8/8 negative controls passing. No negative controls have been run on the physics corpus yet.

Why grade narrators per domain instead of globally?

A model precise on historical dates may be unreliable on quantum-mechanical equations. The registry key is (narrator, domain), not narrator alone. This mirrors classical practice — rijāl scholars graded transmitters domain-specifically as a matter of course.

What happens when a model version changes?

A model-version bump resets the narrator to ungraded until re-evaluated. Version drift is treated as a new narrator — track record does not carry forward. This is intentional: a model at v2 may behave differently from the same model at v1, and inherited reputation would be misleading. Known blind spot: hosted models whose weights change without a version-string change will silently retain a grade they no longer merit.

What is the "unverifiable" verdict and why does it matter?

"Unverifiable" is the third column of the decision matrix — the classical position of tawaqquf (suspension of judgment). A content critic that cannot evaluate a claim must be able to say so rather than defaulting to "consistent." Routing unverifiable claims conservatively is what makes the framework safe under a weak critic. It is also the binding constraint on practical coverage — most real claims receive this verdict under the reference critic.

What did the evaluation fail to show?

Three things, reported in the same typographic weight as the successes: (1) The jarḥ–taʿdīl loop recovered 3 of 4 narrator grades but missed the highest-fault narrator through insufficient calibration evidence. (2) The confidence baseline is synthetic and proves nothing about real LLM self-confidence. (3) The matn critic's false-consistent rate on meaning-changing corruption is 39.1% (#126) — it is a rejection gate, not an acceptance gate. See the evaluation section.

Is this a religious claim or ruling?

No. ISNAD borrows a methodology — the epistemological architecture of isnād criticism — from classical Islamic hadith science. It makes no religious claims, issues no rulings, and does not position any computational system as an authority within the religious sciences. See the positionality statement.

Can I use this in production today?

The offline pipeline has been exercised end to end on 20,000 claims, and the end-to-end utility experiment (#128) runs live LLM-generated claims through the real pipeline — gated serving at 0.0% vs 50.0% ungated served-error. The LangChain integration is shipped and tested. A live production deployment where chains are written automatically at ingest and the registry gates a serving path under real query load is still open work.

How do I cite it?

Two BibTeX entries are available in the Cite section — one for the arXiv paper and one for the software. Paper: arXiv:2607.24117 (cs.AI). Software DOI: 10.5281/zenodo.21216872.

What licence is it under?

Code: Apache-2.0 — with a public commitment that src/isnad/ stays Apache-2.0 in perpetuity (see LICENSING.md). Paper and documentation: CC BY 4.0.

How can I contribute?

The CONTRIBUTING.md lists good first issues — alternative critics, seed-grade bootstrappers, correlation detectors, and pipeline adapters. The highest-impact open problem is chain independence (#54). Pull requests welcome.