إسناد

ISNAD

Claim-level trust for multi-agent AI systems.

Every claim carries its chain. Every transmitter is graded. Adapted from 1,200 years of hadith transmission science.

arXiv:2607.24117 · cs.AI / cs.MA · 25 pages · pip install isnad v2.0.5 · Apache-2.0
PyPI version GitHub stars License arXiv

The Living Chain#

DESTRUCTIVE ▼ GENERATIVE ▲ GENERATIVE ▲ source RELIABLE pdf- scraper RELIABLE ingest- model UNGRADED answer- model RELIABLE claim CAP Grade ṣaḥīḥ Action SERVE WITH CAVEAT

Every claim carries its chain.

source
RELIABLE
pdf-scraper
RELIABLE
DESTRUCTIVE ▼
ingest-model
UNGRADED
GENERATIVE ▲
answer-model
RELIABLE
GENERATIVE ▲
ḥasan
SERVE WITH CAVEAT
One ungraded narrator caps the grade.
Interactive

Try the Decision Matrix#

Select a chain grade and content verdict — the matrix routes your claim.

Route
SERVE WITH CAVEAT
Show full 4×3 decision matrix
ConsistentContradictionUnverifiable
Ṣaḥīḥ-tierServe directly; cacheʿIlal flag → human reviewServe with caveat
Ḥasan-tierServe with caveat; seek corroborationHold; do not serveHold in review queue
Ḍaʿīf-tierHold; seek corroborationQuarantineHold in review queue
Mawḍūʿ-tierReject & quarantine narratorReject & quarantine narratorReject & quarantine narrator
Interactive

Corroboration & the Madār#

Two chains, same claim. Does corroboration fire — or does the madār block it?

Chain 1 — Base (Ḍaʿīf) src-A scrape-A ingest-A ḍaʿīf-tier ✓ Narrator identity ✓ Model family ✓ Upstream source Chain 2 — Corroborating (Ṣaḥīḥ) src-B scrape-B ingest-B ṣaḥīḥ-tier Independence: 3/3 ✓ CORROBORATION FIRES ḍaʿīf → ḥasan (capped)
Honest annotation: Validated 603/603 on Wikipedia, 104/104 on physics textbooks, 8/8 negative controls. But no negative controls were run on the physics corpus, and both Wikipedia corpora are Wikimedia projects — so v2 validates the mechanics of corroboration, not the independence assumption. Content-level correlated error is not addressed at all.
Section 1

The Problem#

When a multi-agent AI system answers you, that claim has passed through many hands — a source, a scraper, an ingestion model, a synthesis model. Each hand can drop, distort, or invent.

A single factual claim may pass through document retrieval, extraction, summarization, knowledge integration, and answer synthesis before reaching a user. Provenance systems can reconstruct that chain after the fact. They provide limited guidance on the question users actually care about: this specific claim, having passed through these specific transformers — should I trust it?

Provenance tools record what happened. Nothing grades who did it — or tells you how much to trust the result. A confident model at the end of a chain cannot repair a corrupted extraction at the start of it.

Section 2

The Idea#

Classical Islamic hadith science spent twelve centuries on a structurally identical problem: whether to trust knowledge transmitted through chains of human narrators. ISNAD transfers that methodology to AI pipelines.

1

Every claim carries its transmission chain (isnād)

2

Every transmitter keeps a living, per-domain grade (rijāl)

3

The weakest link caps the chain's trust

4

Independent corroboration upgrades — capped and gated (mutābaʿāt)

5

Content is criticized separately from the chain (matn)

Section 3

Concept Mapping: From Ḥadīth Science to AI Pipelines#

The intellectual heart of the framework — the paper's Table 1, mapping each classical term to its ISNAD usage.

TermClassical meaningUsage in ISNAD
isnādThe complete chain of transmitters attached to a narrationThe agent chain attached to a claim
matnThe transmitted content itselfThe claim text
rijāl"The transmitters"; the discipline of evaluating themThe narrator registry
ʿadālahA narrator's integrity / uprightnessResistance to manipulation; source trustworthiness
ḍabṭA narrator's precision and retention accuracyMeasured error rate of an agent/model
al-jarḥ wa-l-taʿdīlFormal criticism and accreditation of narratorsThe registry update process (evals, audits)
ittiṣāl / munqaṭiʿChain continuity / a broken chainComplete trace / missing trace segment
ṣaḥīḥ, ḥasan, ḍaʿīf, mawḍūʿSound, good, weak, fabricated (grades)Trust tiers for claims
mutābaʿātIndependent corroborating chainsMultiple disjoint agent chains asserting one claim
ʿilalHidden defects: sound-looking chain, defective contentChain passes but content criticism fails
muḥaddithThe expert scholar who adjudicates difficult casesThe human reviewer in the loop

Refining the Weakest-Link Rule

In hadith transmission a weak link is unconditionally corrupting. AI transformations are not all alike. A destructive transformation — extraction, chunking, lossy summarization — can only lose or corrupt information, and there the strict minimum applies: nothing downstream recovers what the extractor dropped. A generative transformation can either repair upstream noise or introduce fresh corruption by favouring its parametric prior over retrieved evidence. For generative steps the rule is therefore not a blind minimum but a bounded one: a generative narrator can raise the floor only up to its own grade, and only when corroboration supports the repair — and it can always lower it.

This mirrors a distinction the muḥaddithūn drew themselves, between riwāya bi-l-lafẓ (transmission of exact wording) and riwāya bi-l-maʿnā (transmission by meaning). The weakest-link principle survives; it is refined by transformation type.

Domain-conditioned grading. The registry key is (narrator, domain), not narrator alone. A model precise on historical dates may be unreliable on quantum-mechanical equations. This is a return to classical practice — rijāl scholars graded transmitters domain-specifically as a matter of course.
Version-bump resets. A model-version bump resets a narrator to ungraded until re-evaluated. Version drift is treated as a new narrator, not inherited reputation. Known blind spot: a hosted model whose weights change without a version-string change will silently retain a grade it no longer merits. Deployments against third-party APIs should treat scheduled re-evaluation as mandatory.
Section 4

The Decision Matrix#

The framework's most actionable artifact — chain grade × content verdict → action. The third column is the binding constraint on real-world coverage.

Matn: consistentMatn: contradictionMatn: unverifiable
Ṣaḥīḥ-tier chain
complete; all narrators graded reliable
Serve directly; cacheʿIlal flag → human review. Highest-value review case.Serve with caveat
Ḥasan-tier chain
complete; ≥1 narrator ungraded or mid-tier
Serve with explicit confidence caveat; seek corroborationHold in review queue; do not serveHold in review queue
Ḍaʿīf-tier chain
weak narrator, or munqaṭiʿ
Hold; seek corroborating chain before servingQuarantineHold in review queue
Mawḍūʿ-tier chain
rejected narrator — e.g. known injection source
Reject and log; quarantine the narratorReject and log; quarantine the narratorReject and log; quarantine the narrator
The third column is not a formality. Content criticism is a partial function. A critic that cannot evaluate a claim must be able to say so rather than defaulting to "consistent." Unverifiable is the classical position of tawaqquf — suspension of judgment. Routing it conservatively is what makes the framework safe under a weak critic. It is also the binding constraint on practical coverage. A framework that omitted this column would look far more capable on paper and behave far worse in deployment.
Sound chain + contradiction is the most informative signal in the system, not an error state. It means either a trusted source has changed the world's state, or the corpus contains a latent defect. Both are exactly what a self-maintaining knowledge base exists to catch.
The mawḍūʿ tier is the native defence against knowledge-base poisoning. A compromised source or manipulated agent is, in rijāl terms, a fabricator — a narrator whose ʿadālah has failed. Grading for integrity and not merely precision means the framework quarantines an injection source rather than merely noting low confidence in its output.

Contradictions go to humans by default. Knowledge-conflict research consistently finds LLMs unreliable at representing and reconciling competing evidence, and during prototype development LLM synthesis handled contradiction cases worse than an extractive baseline until contradiction markers were made explicitly visible in context — after which auto-resolution was still left disabled in favour of gated review.

Section 5

Worked Example: End to End#

Claim: "the momentum of a photon is p = h/λ", ingested from an openly licensed physics text.

Chain: [OpenStax University Physics Vol. 3 — source; publisher-trusted, ʿadālah high] → [pdf-scraper v1.2 — ḍabṭ graded: high extraction fidelity on text-layer PDFs] → [ingest-analysis model M@v — ungraded, recent version bump] → [ingest-renderer M@v — ungraded]

Chain grade: complete (ittiṣāl holds), but two narrators ungraded → ḥasan-tier by the weakest-link rule.

Matn criticism: the corpus already contains a page asserting momentum as p = mv without regime qualification. Contradiction detected.

Matrix cell: ḥasan × contradiction → review queue; do not serve.

Adjudication: a human reviewer recognises both claims as valid in different regimes — classical and quantum — and updates both pages with regime qualifiers, resolving the contradiction once, at ingest time, rather than leaving every future query to renegotiate it. The resolution feeds back: the ingestion model's contradiction-detection behaviour on this case becomes taʿdīl evidence toward its grade.

This example is implemented as a runnable demo and an integration test in the reference implementation (make demo, examples/worked_example.py, tests/test_worked_example.py).

Section 6

What ISNAD Adds Over Existing Provenance#

Each Mechanism Against Its Strongest Existing Ancestor

MechanismStrongest ancestorWhat ISNAD adds
Chain aggregationJøsang subjective logic (2001) — trust discounting along chainsTransformation typing (riwāya bi-l-lafẓ vs. bi-l-maʿnā); completeness semantics
Source registryTruth discovery (Yin et al. 2008; Dong et al. 2014)Per-domain grading; version-bump resets; quarantine lifecycle
Content criticismFEVER-style fact-checkingDecoupled from chain quality; operates on matn independently; explicit unverifiable verdict
Agent gradingMAS reputation (EigenTrust, ReGreT, FIRE)Claim-level provenance, not just partner selection
Decision routing— (no direct ancestor)Chain grade × content verdict → concrete serve/review/quarantine action

Capability Coverage

ApproachChain captureGraded transmittersContent criticismDecision routing
W3C PROV / PROV-AGENT
Truth discovery (TruthFinder, KBT)partial
MAS reputation (EigenTrust, ReGreT)partial
FEVER-style fact-checking
ISNAD

Chain capture means claim-level transmission chains are recorded per claim; graded transmitters means each transformer carries a maintained reliability grade; content criticism means content is evaluated independently of chain quality; decision routing means the system emits a concrete serve/review/quarantine action. A ✅ indicates the approach addresses that capability as a first-class concern.

Deep lineage: Condorcet's jury theorem (1785), Hume on testimony (1748), the hearsay-within-hearsay rule (US FRE 805), Dempster–Shafer evidence combination, Dawid & Skene (1979), Jøsang's subjective logic (2001), TruthFinder (2008), Knowledge Vault (2014), Knowledge-Based Trust (2015). The paper's claim is not novelty for its own sake but the transfer of the most operationally complete solution to a domain that lacks it.

Section 7

Quickstart#

Current release: v2.0.5 · Python 3.12+ · Apache-2.0 · 157 tests passing

Install

pip install isnad

60-Second Example

from isnad import Registry, Chain, ChainLinkSpec, grade_chain, decide
from isnad.types import NarratorGrade, ContentVerdict
from isnad.critics import EmbeddingCritic

# Build a chain: source → scraper → model
chain = Chain([
    ChainLinkSpec("openstax-textbook", 0, domain="physics"),
    ChainLinkSpec("pdf-scraper-v2", 1),
    ChainLinkSpec("ingest-model-v3", 2),
])

# Seed-grade known narrators
reg = Registry()
reg.register("openstax-textbook", "physics", grade=NarratorGrade.RELIABLE)
reg.register("pdf-scraper-v2",    "physics", grade=NarratorGrade.RELIABLE)
reg.register("ingest-model-v3",   "physics", grade=NarratorGrade.ACCEPTABLE)

# Grade the chain
grades     = [reg.get_grade(l.narrator_id, l.domain) for l in chain.links]
transforms = [l.transform_type for l in chain.links]
chain_grade = grade_chain(grades, transforms, is_complete=True)

# Content criticism — decoupled from chain grading
critic  = EmbeddingCritic()
verdict = critic.evaluate("p = h/λ", "p = h/lambda", ["p = mv"])
action  = decide(chain_grade, verdict)

print(f"Chain: {chain_grade.value.upper()} | Content: {verdict.value} | Action: {action.value}")

LangChain Integration (5 Lines)

pip install isnad[langchain]

from isnad.integrations.langchain import IsnadTracer, seed_registry
from isnad.critics import EmbeddingCritic

reg = seed_registry({"source:docs": "reliable", "model:gpt-4o": "acceptable"})
tracer = IsnadTracer(registry=reg, critic=EmbeddingCritic())
chain.invoke("What is F=ma?", config={"callbacks": [tracer]})
print(tracer.report())

Run the Experiments

git clone https://github.com/alizahidraja/isnad && cd isnad
make install   # uv sync
make test      # zero config, SQLite fallback
make demo      # the paper's worked example (§4.5)
make check     # lint + type-check + test

No database required for pure-logic tests; PostgreSQL is optional (docker compose up, set ISNAD_DATABASE_URL).

Section 8

Architecture & Module Map#

ConceptWhat it doesModule
isnād (chain)Ordered, gap-checked transmission chain per claimisnad/core/chain.py
rijāl (registry)Graded narrator store per (alias@version, domain)isnad/core/registry.py
jarḥ–taʿdīlEvidence-driven state machine for narrator gradesisnad/core/registry.py
Bayesian gradingBeta-distribution narrator grades (default)isnad/core/registry.py
ittiṣālCompleteness as epistemic property (gap → ḍaʿīf)isnad/core/chain.py
Weakest-link gradingChain grade = refined minimum over narratorsisnad/core/grading.py
mutābaʿāt (corroboration)Independent-chain upgrade + madār detectionisnad/core/corroboration.py
matn criticismContent evaluated independently of chain qualityisnad/critics/
Decision matrixChain × content → action routerisnad/core/decision.py
PersistenceSQLAlchemy-backed registry (swap via protocol)isnad/storage/
APIFastAPI service with DI + Prometheus metricsisnad/api/
CLIisnad serve, isnad seedisnad/cli/
ʿadālah / ḍabṭIntegrity and precision as two distinct axesisnad/types.py

Pluggable Strategies — Updated Defaults

StrategyProtocolDefaultWhat it decides
GradingStrategyisnad/types.pyRefinedWeakestLinkHow link grades combine into a chain grade
TransitionPolicyisnad/types.pyBayesianTransitionPolicyHow evidence moves narrators between states
CorroborationPolicyisnad/types.pyCappedCorroborationPolicyHow independent chains upgrade a claim
CorrelationDetectorisnad/types.pySharedLineageDetectorWhether two chains are truly independent
ContentCriticisnad/critics/base.pyHybridCritic / EmbeddingCriticContent contradiction detection

Reference Schema

CREATE TABLE rijal_claims (
    claim_id         TEXT PRIMARY KEY,
    page_slug        TEXT NOT NULL,
    claim_text       TEXT NOT NULL,
    narrator_chain   JSONB NOT NULL,
    chain_confidence NUMERIC(4,3),
    valid_from       TIMESTAMPTZ,
    valid_until      TIMESTAMPTZ,
    superseded_by    TEXT,
    chain_status     TEXT DEFAULT 'active'
);

CREATE TABLE narrator_registry (
    narrator_id      TEXT NOT NULL,
    domain_tag       TEXT NOT NULL,
    narrator_type    TEXT NOT NULL,
    grade            TEXT NOT NULL,
    known_error_rate NUMERIC(4,3),
    model_version    TEXT,
    is_active        BOOLEAN DEFAULT TRUE,
    PRIMARY KEY (narrator_id, domain_tag)
);

Because the pipeline's transformation points already emit trace identifiers, populating narrator_chain is an instrumentation task, not an architectural change.

Section 9

Evaluation: What's Validated, What Failed, What's Inconclusive#

Setup

  • 20,000 atomic claims extracted from real PDF text: OpenStax University Physics Vols 1–3 (CC BY 4.0) and Crowell's Light and Matter (CC BY-SA), split evenly, 10,000 per source.
  • Domain tags: mechanics (7,600), general (6,890), electromagnetism (3,054), modern-quantum (1,728), optics-waves (728).
  • Extraction via deepseek-chat over chunked PDF text.
  • Split 30% calibration / 70% evaluation → 14,001 evaluation claims, across 10 random seeds.
  • Four narrators with designed fault rates: pdf-scraper@1.2 (1%), ingest@good (2%), ingest@weak (15%), pdf-scraper@0.9-legacy (18%).
  • Faults: OCR noise, digit swap, sign flip, negation drop, unit corruption, formula mangling, entity swap, fabricated numerics, regime confusion. ~5% of chains marked incomplete.
  • Review budgets: B ∈ {2%, 5%, 10%, 20%}.
Definitions. Error: a served claim counts as an error if its text was altered by fault injection, per the manifest — synthetic corruptions of correct source text, not naturally occurring factual errors. Coverage: the fraction of evaluation claims the decision matrix routes to serve or serve-with-caveat. Claims routed to review beyond budget, and quarantined claims, do not count.

Status at a Glance

ComponentStatusWhat remains
Weakest-link grading and quarantine✅ Validated
jarḥ–taʿdīl grade recovery⚠️ PartialRecovered 3/4; missed the highest-fault narrator through insufficient per-cell calibration evidence
Corroboration (mutābaʿāt)✅ ValidatedNegative controls on the physics corpus; organically weak chains
Confidence baseline⛔ Not testedBaseline is synthetic; real model self-confidence untested
Matched-coverage advantage❓ InconclusiveRequires a critic that can reach matched coverage
Content criticism⚠️ PartialSemantic or LLM-backed critic
End-to-end serving pipeline🔲 OpenLive deployment at practical coverage

Validated: Weakest-Link Quarantine

Every claim whose chain contained a narrator graded rejected was quarantined — 4,057 claims, 29% of the evaluation split — and every quarantine was traceable to the specific narrator grade that caused it.

Step 0: source:openstax RELIABLE Step 1: pdf-scraper@1.2 RELIABLE [DESTRUCTIVE] Step 2: ingest@weak REJECTED [GENERATIVE] <- binding link Chain grade: MAWDU -> REJECT_AND_QUARANTINE_NARRATOR

Partial Failure: The jarḥ–taʿdīl Loop

The loop recovered three of four narrator grades from audit evidence alone:

  • pdf-scraper@1.2 (1% fault) → reliable in all 50 (narrator, domain) cells
  • ingest@good (2%) → acceptable in all 50
  • ingest@weak (15%) — deliberately left ungraded — independently discovered and driven to rejected in 49/50 cells, weak in one
The fourth was not recovered, and the failure is the most instructive result in the paper. pdf-scraper@0.9-legacy, carrying the highest designed fault rate (18%), never accumulated enough claims in the calibration split to earn any grade in any domain, and remained ungraded throughout. It was among the largest single sources of injected faults. The loop found the unreliable narrator it was given enough evidence about, and silently missed the one it was not.

Grade recovery is bounded by calibration coverage per (narrator, domain) cell, not only by the transition policy. Domain-conditioned grading multiplies the number of cells that must be filled, and a narrator rare within a domain is invisible to jarḥ–taʿdīl regardless of how unreliable it is. A deployment should monitor per-cell evidence counts as a first-class operational metric and treat a persistently under-evidenced cell as a risk, not a neutral absence.

Validated: Corroboration Across Three Corpora

v1 (exact)v2 (Wikipedia)v3 (physics)
Matching criterionexact stringcosine ≥ 0.75cosine ≥ 0.80
Corpus12 article intros30 article pairs2 textbooks
Sentences21510,54422,372
Candidate pairs136662104
Pairs evaluated136603104
Match density6.3%0.5%
Corroboration fired68/136603/603104/104
Negative controlsnone run8/8none run
Independence basissynthetic chainsseparate editorial communitiesdifferent authors and publishers
Difficultyeasymediumhard
Important caveats. v2 independence is imperfect — English and Simple English Wikipedia have separate editorial communities but share the Wikimedia umbrella. v2 validates the mechanics of corroboration rather than the independence assumption. v3 is the stronger independence test (different authors, publishers, decades, and upstream domains). No negative controls were run on v3 — a 100% fire rate without a discrimination check is weaker evidence than the same rate with one. In both v2 and v3 the weak-tier baseline chain was created by assigning one narrator a weak grade because the sources are high quality. The experiment tests whether corroboration upgrades a weak chain when an independent chain corroborates it — not whether it works on chains that are weak for organic reasons.

The eight negative controls (v2) correctly produced no upgrade: no matching text; shared model family (the madār case); corroborators below the grade gate; a rejected base chain; the ṣaḥīḥ cap; shared upstream source; insufficient independent chains; empty corpus.

Inconclusive: The Matched-Coverage Comparison

This is the result most projects would hide. Lead with it.

Comparing served-error rates across conditions operating at different coverage levels is not meaningful — a system serving 5% of claims will trivially show lower error than one serving 100%. The standard remedy is a matched-coverage sweep. It was specified in advance, attempted, and could not be completed.

With the reference content critic, ISNAD-gated serving could not be driven above 4.8% coverage at any review budget between 5% and 50%. The risk–coverage relationship is a single point, not a curve.

Target coverageISNAD coverageISNAD errorBaseline coverageBaseline error
20%4.8% (not reached)0.0%20.0%15.1%
30%4.8% (not reached)0.0%30.0%15.8%
50%4.8% (not reached)0.0%50.0%16.3%
70%4.8% (not reached)0.0%70.0%16.0%
90%4.8% (not reached)0.0%90.0%16.1%

The rows are reported to show what was attempted and where it failed, not as an advantage.

This experiment establishes that ISNAD-gated serving admits no corrupted claims among those it serves. It does not establish that ISNAD outperforms the baseline at matched coverage, because it cannot reach matched coverage with the reference critic. The binding constraint is the content critic, not the grading mechanism. Any future claim of a risk–coverage advantage requires a critic capable of returning a consistent verdict on real prose, and this paper makes no such claim.

The Confidence Baseline Is Uninformative by Construction

A confidence-gated condition was included to represent current practice. It performs at chance — but for a reason that sharply limits what can be concluded from it. The corpus's model_confidence field is computed before fault injection and never updated. Confidence is statistically independent of corruption by construction. The point-biserial correlation is r = 0.04; mean confidence on corrupted claims (0.853) and clean claims (0.848) differ by less than one hundredth.

This is a floor, not a finding. It establishes that the ISNAD condition is not being measured against a baseline that trivially wins, and nothing more. No claim is made about whether real LLM self-confidence carries a defect signal.

Honest Negatives: Transition Policy & Content Critic

Sweeping the downgrade threshold shows a direct tradeoff with no sweet spot:

Downgrade thresholdCoverageServed errorReview precisionGrades recovered
37.1%0.0%0.112/4
69.4%0.0%0.152/4
1010.0%0.0%0.111/4
1510.0%0.0%0.130/4
2510.0%0.0%0.130/4

There is no setting in this range that delivers both coverage and grade recovery. More audit evidence did not help — increasing the per-cell audit budget reduced coverage from 10.0% to 5.9% at 40 and above. The default policy over-penalises reliable narrators as evidence accumulates on small samples.

The content critic is the binding constraint. The reference critic scores word overlap with a negation heuristic rather than using semantic embeddings — it catches only obvious contradictions and returns unverifiable on most real prose. Under the decision matrix, a ḥasan-tier chain meeting an unverifiable verdict routes to review rather than serve. With most claims receiving that verdict, ḥasan-tier claims cannot graduate to serving, and coverage is capped at approximately the review budget regardless of how well chain grading performs. This single fact explains both the 4.8% coverage ceiling and the saturation in the sweep table.

A separate evaluation of the reference critics on 20 labelled claims returned perfect precision and recall — but those contradictions were template-injected and built from patterns the critics already encode. The numbers overstate real capability and are not reported as evidence.

A semantic or LLM-backed critic is the single highest-value component for practical end-to-end coverage, and building one is the most immediate item of future work.

Case Study: Contradiction Detection on Real Texts

A prototype compiled-wiki system ingested OpenStax University Physics Vols 1–3 and Crowell's Light and Matter via chunked PDF extraction, with DeepSeek-chat as the ingestion model, single run.

Result: 19 contradictions surfaced across the corpus, each manually reviewed and confirmed genuine — among them classical momentum (p = mv) versus photon momentum (p = h/λ); wave versus particle treatments of light; Newtonian versus relativistic kinematics.

All 19 were Type B — regime distinctions where both claims are true in their domains and resolution means qualification — not Type A errors. This validates the detection substrate and routing path; it does not demonstrate error-correction capability.

Four important caveats: (1) Not reproducible — the prototype is a proprietary system, not part of the open-source release. (2) Precision, not recall — 19/19 confirmed genuine, zero false positives. "19 contradictions found" means "at least 19 exist and were caught," not "exactly 19 exist." (3) All Type B — regime distinctions, not errors. (4) Parametric knowledge confound — the ingestion model already knew undergraduate physics. Every contradiction was confirmed to correspond to explicit statements in the ingested chunks, but the experiment does not isolate retrieved evidence from parametric knowledge.
Section 10

Limitations#

The analogy has edges.

Classical narrators were moral agents whose ʿadālah was a judgment of character formed over a lifetime; models are stochastic artifacts whose "integrity" is an engineering property of their deployment. The mapping is functional, not metaphysical. Classical grading also rested on scholarly consensus formed over generations; a registry populated by one team's evaluations is a far thinner epistemic base, and its grades should be presented with corresponding humility.

Registry cold start.

Weakest-link grading with an empty registry caps everything at ḥasan-tier or below. Arguably correct — new pipelines should earn trust — but it front-loads human review cost. Practical value depends on the jarḥ–taʿdīl loop converging faster than review budgets exhaust. §8 showed the intended mitigation (seed grading) goes less far than anticipated.

Grading humans.

Where human contributors are narrators, the registry becomes a reputation system over people, with the fairness, contestability, and workplace-power questions that implies. Flagged as requiring governance design — contestable grades, transparent criteria — not resolved.

False precision.

chain_confidence as a decimal invites over-trust. The ordinal-first principle is the mitigation; implementations that surface raw numbers to end users without tier context would be misusing the framework.

Independence is an idealisation.

Two chains may share no explicit narrator yet still fail together — if both terminate in the same answer model, or in models from a family with correlated training data. Classical scholars faced the same problem in the madār (the pivot narrator through which ostensibly independent chains converge). The reference implementation enforces an independence check on narrator identity, model family, and upstream source — but a genuinely correlated pair sharing none of those three would still pass, and content-level correlated error (two independent sources repeating the same received mistake) is not addressed at all.

Probabilistic and provisional claims.

The framework assumes claims with determinate truth values. "The Higgs boson mass is 125.1 GeV" carries an implicit uncertainty interval. Contradiction detection between a point estimate and an interval is not defined. The lifecycle columns provide a mechanism for supersession but not a semantics for uncertainty.

Storage and query overhead is estimated, not measured.

Chains are short (typically 3–5 links) so per-claim JSONB is small and bounded; the registry grows with pipeline changes, not corpus size. At the targeted scale — thousands to low millions of claims — neither table is expected to bottleneck. These are design expectations, not measured results. No latency or storage benchmark was run, and none is claimed.

Scope exclusions: Byzantine collusion among agents, cryptographic attestation of traces (complementary, not competing), and open-ended or creative generation. ISNAD grades claims that can in principle be true or false and checked against a corpus. It is a framework for knowledge accumulation, not for tasks where there is no fact of the matter to transmit.

Positionality. I am a Muslim practitioner and builder; Islam & AI, a Quranic NLP platform I founded, is where I first worked seriously with these scholarly traditions. This work adapts a methodology developed by hadith scholars; it makes no religious claims, issues no rulings, and does not position any computational system as an authority within the religious sciences themselves. The direction of intellectual credit should be unambiguous: the framework's rigor belongs to twelve centuries of muḥaddithūn; the transfer to AI systems is the contribution claimed here. I would welcome collaboration with hadith-studies scholars to correct any misrepresentation of the classical discipline.
Section 11

What's Next#

Shipped Since the Paper's Submission Version

  • Bayesian (Beta-distribution) narrator grading as the default transition policy, replacing hardcoded thresholds; ISNAD_POLICY env override
  • Working content critics: EmbeddingCritic (TF-IDF, offline), HybridCritic (NLI), LLMCritic
  • Seed-grade bootstrapping — validated as improving coverage from ~5% to ~10%; ISNAD_SEED_CONFIG env var
  • LangChain integration: IsnadTracer callback handler + seed_registry, 9 integration tests passing
  • FastAPI service with dependency injection and Prometheus metrics; CLI
  • SQLAlchemy-backed persistence behind a swappable storage protocol
  • Endpoint identity / model-version-drift handling (alias@version keying)
  • Corroboration v3 on physics textbooks (104/104)
  • Architecture diagram (docs/ARCHITECTURE.drawio)

Highest-Priority Open Work (collaborators welcome)

  1. A semantic or LLM-backed content critic. The single highest-value component for practical coverage — the binding constraint on every headline result.
  2. A live production deployment where chains are written automatically at ingest and the registry gates a serving path under real query load.
  3. Negative controls on the physics corroboration corpus (v3) and corroboration tested on organically weak chains.
  4. Repeated-narrator penalties. Two passes through the same lossy transformation are strictly worse than one. Recorded but not yet penalised.
  5. Per-cell evidence-count monitoring as a first-class operational metric.
  6. Extension to probabilistic and provisional claims.

Good First Issues

  • An alternative critic (sentence-transformers embedding, CrewAI integration)
  • A seed-grade bootstrapper from published benchmark accuracies
  • A new CorrelationDetector using embedding similarity or model-card lineage
  • Extend semantic corroboration to multi-source corpora
  • A calibrated TransitionPolicy from your own pipeline's §8 experiment data
  • Domain-specific ContentCritic with formula canonicalisation
  • Pipeline adapters for CrewAI or AutoGen tracing
Section 12

In the Wild#

Live Verify — Independent Comparison

Paul Hammant — co-creator of Selenium (~22 years ago), maintainer of trunkbaseddevelopment.com, ex-ThoughtWorks, based in Scotland. Reached out to Ali directly after finding the paper.

A dedicated comparison document: Live Verify vs. Isnād–Rijāl

Both attack the same root problemwhy should I trust this claim that reached me through a chain of hands? — and both reach for a provenance answer rather than a plausibility answer. But they sit at opposite ends of the trust pipeline. They are complementary rather than competing.

ISNADLive Verify
LayerInside an AI pipeline, per-claimBetween organisations, per-artifact
MechanismGraded reputation registry + Bayesian loopCryptographic hash + domain endpoint
Trust rootAccumulated narrator track recordAn authorizedBy chain to a jurisdiction's root namespace
Outputserve / review / quarantineverified / not-verified + status
Truth stanceProbabilistic (how reliable)Binary (unaltered vs. altered)
HandlesDistortion by AI transformersForgery/tampering of documents
Both frameworks are chain-based, but the chains carry different things. ISNAD's isnād is a transmission chain — the ordered list of hands a claim passed through, each graded on how reliably it forwards claims. Live Verify's authority chain is a vouching chain — not who relayed the artifact, but who vouches the issuer is entitled to issue it. So where ISNAD asks "how trustworthy is this path of transmitters?", Live Verify asks "is this issuer a legitimate authority?" Both refuse to stop at the leaf; they just walk in different directions — ISNAD backwards through transmission, Live Verify upwards through authorization.
The composition argument. A Live Verify seal is an ideal high-trust narrator input to an ISNAD chain. A cryptographically-anchored, authority-chained source is exactly the kind of link that deserves a top narrator grade — its integrity axis (ʿadālah) is anchored by cryptography and by an authority chain to a sovereign root, rather than by accumulated track record. A Live Verify seal can bootstrap a narrator to a high grade on day one, before any evidence history exists.
Verify this page's own verdict with your phone. ISNAD now issues its own sealed verdicts as a Live Verify issuer. Select the verdict text at alizahidraja.com/verify/isnad-demo and run a Live Verify client — it resolves to verified / self-verified (amber), not green. A page about provenance that carries its own verifiable provenance is the demo. See the issuer surface →

"ISNAD's README carries a 'What's Validated vs. What's Not' box and reports two of its own results as inconclusive rather than positive. … Both projects treat scoping-out over-claims as a first-class feature, not an afterthought." — from the Live Verify comparison document.

Read the full comparison → · Have your own project or comparison? Add it — reach out.

Indexing, Archiving & Aggregation

Also discoverable via NASA ADS, Google Scholar, Semantic Scholar, Connected Papers, alphaXiv, CatalyzeX, Litmaps, scite, OpenAIRE, and Hugging Face Daily Papers. Reddit launch threads — URLs pending; will be linked when recovered.

Prior & Adjacent Work

  • PROV-AGENT (Souza et al., IEEE e-Science 2025) — extends W3C PROV to capture agent interactions.
  • Agent-Sentry (Sequeira et al., 2026) — bounding LLM agents via execution provenance.
  • HadithRank (Hawramani) — models isnād topology using information theory, treating independent chains as redundant channels. That a mathematician grading hadith and this paper grading agents converge on the same redundancy rule is evidence the mapping is structural, not decorative.
  • The Provenance Paradox (Prakash, 2026) — when delegates can inflate self-reported quality, quality-based routing can perform worse than random. Consistent with ISNAD's design intuition but not a replication.
  • RAPS (2026) — reputation-aware publish–subscribe with Bayesian overlays.
  • TrustTrade (2026) — replaces "uniform trust" with cross-agent consistency weighting.
  • AgentPoison (Chen et al., NeurIPS 2024) — red-teaming LLM agents by poisoning memory/knowledge bases; the demonstrated threat the mawḍūʿ tier defends against.
  • Karpathy's "LLM Wiki" sketch (2026) — the compiled-knowledge-base system class that motivates the whole framework.

Two public repositories independently reference isnad-style provenance chains: cognalith/isnad (agent skill security with provenance verification) and VAMaksimov/threat-intel-format (threat intelligence schema with isnad-style provenance chains). Neither is affiliated.

A note on the name. "ISNAD" here refers to the Isnād–Rijāl Framework described in Grading the Narrators (arXiv:2607.24117). Several unrelated projects in the AI-trust space use the same Arabic word — it is a common noun meaning "chain of transmission" and no one owns it. This project is open-source under Apache-2.0, has no token, no blockchain component, and no commercial affiliation with any similarly named product.
Section 14

Cite#

Paper

@article{raja2026grading,
  author        = {Ali Zahid Raja},
  title         = {Grading the Narrators: An Isnad--Rijal Framework for
                   Claim-Level Provenance in Multi-Agent Knowledge Systems},
  year          = 2026,
  doi           = {10.48550/arXiv.2607.24117},
  eprint        = {2607.24117},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI}
}

Software

@software{raja2026isnad,
  author = {Ali Zahid Raja},
  title  = {Isnad--Rijal Framework: Reference Implementation},
  year   = 2026,
  doi    = {10.5281/zenodo.21216873},
  orcid  = {0009-0003-7875-4590}
}

LLM usage disclosure: drafting and editing were assisted by a large language model; all technical claims, design decisions, experimental results, and final text were reviewed by and are the responsibility of the author.

FAQ

Frequently Asked Questions#

What is ISNAD in one sentence?

ISNAD is an open-source framework that grades every agent, scraper, and model in a claim's transmission chain — adapted from classical Islamic hadith science — to tell you how much to trust a claim that has passed through multiple AI hands.

How is this different from W3C PROV or PROV-AGENT?

PROV records what happened — it captures chains, agents, and activities. ISNAD grades who did it and whether to trust the result. PROV gives you a transmission log; ISNAD tells you which link in that log is the weakest, and what action to take. See the positioning tables for a detailed comparison.

Does ISNAD outperform confidence-based gating?

The matched-coverage comparison could not be completed. With the reference content critic, ISNAD-gated serving could not be driven above 4.8% coverage, so a meaningful head-to-head comparison was not possible. ISNAD admitted zero corrupted claims among those it served — but it served very few. See the full evaluation.

What does "the chain is graded by its weakest link" mean in practice?

A ṣaḥīḥ model at the end of a chain cannot rescue a ḍaʿīf scraper at the start. The chain's grade is the minimum grade among its narrators — adjusted by transformation type (destructive vs. generative) — because information lost early in the chain cannot be recovered downstream. The chain animation demonstrates this visually.

What is mutābaʿāt / corroboration, and when does it not fire?

Mutābaʿāt is the mechanism by which independent chains asserting the same claim corroborate each other, upgrading the claim's grade (capped at ṣaḥīḥ). It does not fire when chains share a narrator, model family, or upstream source — the classical madār case. It is validated at 603/603 on Wikipedia and 104/104 on physics textbooks, with 8/8 negative controls passing. No negative controls have been run on the physics corpus yet.

Why grade narrators per domain instead of globally?

A model precise on historical dates may be unreliable on quantum-mechanical equations. The registry key is (narrator, domain), not narrator alone. This mirrors classical practice — rijāl scholars graded transmitters domain-specifically as a matter of course.

What happens when a model version changes?

A model-version bump resets the narrator to ungraded until re-evaluated. Version drift is treated as a new narrator — track record does not carry forward. This is intentional: a model at v2 may behave differently from the same model at v1, and inherited reputation would be misleading. Known blind spot: hosted models whose weights change without a version-string change will silently retain a grade they no longer merit.

What is the "unverifiable" verdict and why does it matter?

"Unverifiable" is the third column of the decision matrix — the classical position of tawaqquf (suspension of judgment). A content critic that cannot evaluate a claim must be able to say so rather than defaulting to "consistent." Routing unverifiable claims conservatively is what makes the framework safe under a weak critic. It is also the binding constraint on practical coverage — most real claims receive this verdict under the reference critic.

What did the evaluation fail to show?

Three things: (1) ISNAD cannot outperform a baseline at matched coverage because it cannot reach matched coverage — the content critic caps it at 4.8%. (2) The jarḥ–taʿdīl loop recovered 3 of 4 narrator grades but missed the highest-fault narrator because it lacked sufficient calibration evidence. (3) The confidence baseline is synthetic and proves nothing about real LLM self-confidence. These are reported in the same typographic weight as the successes. See the evaluation section.

Is this a religious claim or ruling?

No. ISNAD borrows a methodology — the epistemological architecture of isnād criticism — from classical Islamic hadith science. It makes no religious claims, issues no rulings, and does not position any computational system as an authority within the religious sciences. See the positionality statement.

Can I use this in production today?

The offline pipeline — chain grading, narrator registry, content criticism, decision routing, corroboration — has been exercised end to end on 20,000 claims. The LangChain integration is shipped and tested. A live production deployment where chains are written automatically at ingest and the registry gates a serving path under real query load is open work — it is the highest-priority next step and collaborators are welcome.

How do I cite it?

Two BibTeX entries are available in the Cite section — one for the arXiv paper and one for the software. Paper: arXiv:2607.24117 (cs.AI). Software DOI: 10.5281/zenodo.21216873.

What licence is it under?

Code: Apache-2.0. Paper and documentation: CC BY 4.0. Both are permissive open-source/open-access licences suitable for academic and commercial use.

How can I contribute?

The CONTRIBUTING.md lists good first issues — including alternative critics, seed-grade bootstrappers, new correlation detectors, and pipeline adapters. The highest-impact contribution is a semantic or LLM-backed content critic, which is the binding constraint on practical coverage. Pull requests welcome.