ISNAD
Claim-level trust for multi-agent AI systems.
Every claim carries its chain. Every transmitter is graded. Adapted from 1,200 years of hadith transmission science.
ISNAD in production — audit export#
The reference implementation now emits a tamper-evident AuditRecord for every graded claim — the evidence artifact auditors and regulators will ask for. ISNAD produces evidence artifacts; it does not confer conformity with any regulation.
isnad export --claim <claim-id> --format json --verify
Truncated example record (illustrative):
{
"record_id": "audit-…",
"record_version": "1",
"claim_id": "c1",
"claim_text": "p = mv",
"final_grade": "hasan",
"chain": [
{ "narrator_id": "src", "narrator_type": "dataset", "grade": "reliable" },
{ "narrator_id": "m", "narrator_type": "model", "grade": "acceptable" }
],
"weakest_link": { "narrator_id": "m", "why": "lowest grade" },
"integrity": { "record_hash": "sha256:…", "hash_algorithm": "sha256" }
}
The record hash is SHA-256 over RFC 8785-canonical JSON. Verify it with isnad verify --record record.json. The full field mapping — which AuditRecord field supports which requirement under the EU AI Act, ISO/IEC 42001, NIST AI RMF, and SDAIA — is at docs/evidence-mapping.md.
RELIABLE narrator grade is an operator-assigned judgment, not a verified fact about the upstream source. See the evidence architecture offer →
The Living Chain#
Every claim carries its chain.
Try the Decision Matrix#
Select a chain grade and content verdict — the matrix routes your claim.
Show full 4×3 decision matrix
| Consistent | Contradiction | Unverifiable | |
|---|---|---|---|
| Ṣaḥīḥ-tier | Serve directly; cache | ʿIlal flag → human review | Serve with caveat |
| Ḥasan-tier | Serve with caveat; seek corroboration | Hold; do not serve | Hold in review queue |
| Ḍaʿīf-tier | Hold; seek corroboration | Quarantine | Hold in review queue |
| Mawḍūʿ-tier | Reject & quarantine narrator | Reject & quarantine narrator | Reject & quarantine narrator |
Corroboration & the Madār#
Two chains, same claim. Does corroboration fire — or does the madār block it?
The Problem#
When a multi-agent AI system answers you, that claim has passed through many hands — a source, a scraper, an ingestion model, a synthesis model. Each hand can drop, distort, or invent.
A single factual claim may pass through document retrieval, extraction, summarization, knowledge integration, and answer synthesis before reaching a user. Provenance systems can reconstruct that chain after the fact. They provide limited guidance on the question users actually care about: this specific claim, having passed through these specific transformers — should I trust it?
Provenance tools record what happened. Nothing grades who did it — or tells you how much to trust the result. A confident model at the end of a chain cannot repair a corrupted extraction at the start of it.
The Idea#
Classical Islamic hadith science spent twelve centuries on a structurally identical problem: whether to trust knowledge transmitted through chains of human narrators. ISNAD transfers that methodology to AI pipelines.
Every claim carries its transmission chain (isnād)
Every transmitter keeps a living, per-domain grade (rijāl)
The weakest link caps the chain's trust
Independent corroboration upgrades — capped and gated (mutābaʿāt)
Content is criticized separately from the chain (matn)
Concept Mapping: From Ḥadīth Science to AI Pipelines#
The intellectual heart of the framework — the paper's Table 1, mapping each classical term to its ISNAD usage.
| Term | Classical meaning | Usage in ISNAD |
|---|---|---|
| isnād | The complete chain of transmitters attached to a narration | The agent chain attached to a claim |
| matn | The transmitted content itself | The claim text |
| rijāl | "The transmitters"; the discipline of evaluating them | The narrator registry |
| ʿadālah | A narrator's integrity / uprightness | Resistance to manipulation; source trustworthiness |
| ḍabṭ | A narrator's precision and retention accuracy | Measured error rate of an agent/model |
| al-jarḥ wa-l-taʿdīl | Formal criticism and accreditation of narrators | The registry update process (evals, audits) |
| ittiṣāl / munqaṭiʿ | Chain continuity / a broken chain | Complete trace / missing trace segment |
| ṣaḥīḥ, ḥasan, ḍaʿīf, mawḍūʿ | Sound, good, weak, fabricated (grades) | Trust tiers for claims |
| mutābaʿāt | Independent corroborating chains | Multiple disjoint agent chains asserting one claim |
| ʿilal | Hidden defects: sound-looking chain, defective content | Chain passes but content criticism fails |
| muḥaddith | The expert scholar who adjudicates difficult cases | The human reviewer in the loop |
Refining the Weakest-Link Rule
In hadith transmission a weak link is unconditionally corrupting. AI transformations are not all alike. A destructive transformation — extraction, chunking, lossy summarization — can only lose or corrupt information, and there the strict minimum applies: nothing downstream recovers what the extractor dropped. A generative transformation can either repair upstream noise or introduce fresh corruption by favouring its parametric prior over retrieved evidence. For generative steps the rule is therefore not a blind minimum but a bounded one: a generative narrator can raise the floor only up to its own grade, and only when corroboration supports the repair — and it can always lower it.
This mirrors a distinction the muḥaddithūn drew themselves, between riwāya bi-l-lafẓ (transmission of exact wording) and riwāya bi-l-maʿnā (transmission by meaning). The weakest-link principle survives; it is refined by transformation type.
(narrator, domain), not narrator alone. A model precise on historical dates may be unreliable on quantum-mechanical equations. This is a return to classical practice — rijāl scholars graded transmitters domain-specifically as a matter of course.
The Decision Matrix#
The framework's most actionable artifact — chain grade × content verdict → action. The third column is the binding constraint on real-world coverage.
| Matn: consistent | Matn: contradiction | Matn: unverifiable | |
|---|---|---|---|
| Ṣaḥīḥ-tier chain complete; all narrators graded reliable | Serve directly; cache | ʿIlal flag → human review. Highest-value review case. | Serve with caveat |
| Ḥasan-tier chain complete; ≥1 narrator ungraded or mid-tier | Serve with explicit confidence caveat; seek corroboration | Hold in review queue; do not serve | Hold in review queue |
| Ḍaʿīf-tier chain weak narrator, or munqaṭiʿ | Hold; seek corroborating chain before serving | Quarantine | Hold in review queue |
| Mawḍūʿ-tier chain rejected narrator — e.g. known injection source | Reject and log; quarantine the narrator | Reject and log; quarantine the narrator | Reject and log; quarantine the narrator |
Contradictions go to humans by default. Knowledge-conflict research consistently finds LLMs unreliable at representing and reconciling competing evidence, and during prototype development LLM synthesis handled contradiction cases worse than an extractive baseline until contradiction markers were made explicitly visible in context — after which auto-resolution was still left disabled in favour of gated review.
Worked Example: End to End#
Claim: "the momentum of a photon is p = h/λ", ingested from an openly licensed physics text.
Chain: [OpenStax University Physics Vol. 3 — source; publisher-trusted, ʿadālah high] → [pdf-scraper v1.2 — ḍabṭ graded: high extraction fidelity on text-layer PDFs] → [ingest-analysis model M@v — ungraded, recent version bump] → [ingest-renderer M@v — ungraded]
Chain grade: complete (ittiṣāl holds), but two narrators ungraded → ḥasan-tier by the weakest-link rule.
Matn criticism: the corpus already contains a page asserting momentum as p = mv without regime qualification. Contradiction detected.
Matrix cell: ḥasan × contradiction → review queue; do not serve.
Adjudication: a human reviewer recognises both claims as valid in different regimes — classical and quantum — and updates both pages with regime qualifiers, resolving the contradiction once, at ingest time, rather than leaving every future query to renegotiate it. The resolution feeds back: the ingestion model's contradiction-detection behaviour on this case becomes taʿdīl evidence toward its grade.
This example is implemented as a runnable demo and an integration test in the reference implementation (make demo, examples/worked_example.py, tests/test_worked_example.py).
What ISNAD Adds Over Existing Provenance#
Each Mechanism Against Its Strongest Existing Ancestor
| Mechanism | Strongest ancestor | What ISNAD adds |
|---|---|---|
| Chain aggregation | Jøsang subjective logic (2001) — trust discounting along chains | Transformation typing (riwāya bi-l-lafẓ vs. bi-l-maʿnā); completeness semantics |
| Source registry | Truth discovery (Yin et al. 2008; Dong et al. 2014) | Per-domain grading; version-bump resets; quarantine lifecycle |
| Content criticism | FEVER-style fact-checking | Decoupled from chain quality; operates on matn independently; explicit unverifiable verdict |
| Agent grading | MAS reputation (EigenTrust, ReGreT, FIRE) | Claim-level provenance, not just partner selection |
| Decision routing | — (no direct ancestor) | Chain grade × content verdict → concrete serve/review/quarantine action |
Capability Coverage
| Approach | Chain capture | Graded transmitters | Content criticism | Decision routing |
|---|---|---|---|---|
| W3C PROV / PROV-AGENT | ✅ | ❌ | ❌ | ❌ |
| Truth discovery (TruthFinder, KBT) | ❌ | ✅ | partial | ❌ |
| MAS reputation (EigenTrust, ReGreT) | ❌ | ✅ | ❌ | partial |
| FEVER-style fact-checking | ❌ | ❌ | ✅ | ❌ |
| ISNAD | ✅ | ✅ | ✅ | ✅ |
Chain capture means claim-level transmission chains are recorded per claim; graded transmitters means each transformer carries a maintained reliability grade; content criticism means content is evaluated independently of chain quality; decision routing means the system emits a concrete serve/review/quarantine action. A ✅ indicates the approach addresses that capability as a first-class concern.
Deep lineage: Condorcet's jury theorem (1785), Hume on testimony (1748), the hearsay-within-hearsay rule (US FRE 805), Dempster–Shafer evidence combination, Dawid & Skene (1979), Jøsang's subjective logic (2001), TruthFinder (2008), Knowledge Vault (2014), Knowledge-Based Trust (2015). The paper's claim is not novelty for its own sake but the transfer of the most operationally complete solution to a domain that lacks it.
Quickstart#
Install
pip install isnad
npm i isnad
Also on npm: the JS/TS audit-record verifier — verifies SHA-256 / Merkle / detached-signature integrity of ISNAD audit records. Grading stays in the Python core.
60-Second Example
from isnad import Registry, Chain, ChainLinkSpec, grade_chain, decide
from isnad.types import NarratorGrade, ContentVerdict
from isnad.critics import EmbeddingCritic
# Build a chain: source → scraper → model
chain = Chain([
ChainLinkSpec("openstax-textbook", 0, domain="physics"),
ChainLinkSpec("pdf-scraper-v2", 1),
ChainLinkSpec("ingest-model-v3", 2),
])
# Seed-grade known narrators
reg = Registry()
reg.register("openstax-textbook", "physics", grade=NarratorGrade.RELIABLE)
reg.register("pdf-scraper-v2", "physics", grade=NarratorGrade.RELIABLE)
reg.register("ingest-model-v3", "physics", grade=NarratorGrade.ACCEPTABLE)
# Grade the chain
grades = [reg.get_grade(l.narrator_id, l.domain) for l in chain.links]
transforms = [l.transform_type for l in chain.links]
chain_grade = grade_chain(grades, transforms, is_complete=True)
# Content criticism — decoupled from chain grading
critic = EmbeddingCritic()
verdict = critic.evaluate("p = h/λ", "p = h/lambda", ["p = mv"])
action = decide(chain_grade, verdict)
print(f"Chain: {chain_grade.value.upper()} | Content: {verdict.value} | Action: {action.value}")
LangChain Integration (5 Lines)
pip install isnad[langchain]
from isnad.integrations.langchain import IsnadTracer, seed_registry
from isnad.critics import EmbeddingCritic
reg = seed_registry({"source:docs": "reliable", "model:gpt-4o": "acceptable"})
tracer = IsnadTracer(registry=reg, critic=EmbeddingCritic())
chain.invoke("What is F=ma?", config={"callbacks": [tracer]})
print(tracer.report())
MCP Integration — Grade MCP Servers as Narrators
from isnad.integrations.mcp import MCPToolObserver, build_mcp_tools, handle_grade_claim
obs = MCPToolObserver(registry, domain="finance")
obs.record_call(tool_name="query", server_name="data-warehouse")
result = obs.grade("revenue grew 12% YoY")
print(result["chain_grade"]) # weakest-link grade over the tool narrators
The Model Context Protocol treats every tool as a transmitter of claims. ISNAD grades MCP servers as narrators — tool outputs are destructive links, so tools stay majhūl (UNGRADED → capped at ḍaʿīf) until the operator seeds or logs evidence. A grade_claim tool exposes the operator's registry to agents at runtime — grades only, never manufactured.
Run the Experiments
git clone https://github.com/alizahidraja/isnad && cd isnad make install # uv sync make test # zero config, SQLite fallback make demo # the paper's worked example (§4.5) make check # lint + type-check + test
No database required for pure-logic tests; PostgreSQL is optional (docker compose up, set ISNAD_DATABASE_URL).
Architecture & Module Map#
| Concept | What it does | Module |
|---|---|---|
| isnād (chain) | Ordered, gap-checked transmission chain per claim | isnad/core/chain.py |
| rijāl (registry) | Graded narrator store per (alias@version, domain) | isnad/core/registry.py |
| jarḥ–taʿdīl | Evidence-driven state machine for narrator grades | isnad/core/registry.py |
| Bayesian grading | Beta-distribution narrator grades (default) | isnad/core/registry.py |
| ittiṣāl | Completeness as epistemic property (gap → ḍaʿīf) | isnad/core/chain.py |
| Weakest-link grading | Chain grade = refined minimum over narrators | isnad/core/grading.py |
| mutābaʿāt (corroboration) | Independent-chain upgrade + madār detection | isnad/core/corroboration.py |
| matn criticism | Content evaluated independently of chain quality | isnad/critics/ |
| Decision matrix | Chain × content → action router | isnad/core/decision.py |
| Persistence | SQLAlchemy-backed registry (swap via protocol) | isnad/storage/ |
| API | FastAPI service with DI + Prometheus metrics | isnad/api/ |
| CLI | isnad serve, isnad seed | isnad/cli/ |
| ʿadālah / ḍabṭ | Integrity and precision as two distinct axes | isnad/types.py |
Pluggable Strategies — Updated Defaults
| Strategy | Protocol | Default | What it decides |
|---|---|---|---|
GradingStrategy | isnad/types.py | RefinedWeakestLink | How link grades combine into a chain grade |
TransitionPolicy | isnad/types.py | BayesianTransitionPolicy | How evidence moves narrators between states |
CorroborationPolicy | isnad/types.py | CappedCorroborationPolicy | How independent chains upgrade a claim |
CorrelationDetector | isnad/types.py | SharedLineageDetector | Whether two chains are truly independent |
ContentCritic | isnad/critics/base.py | HybridCritic / EmbeddingCritic | Content contradiction detection |
Reference Schema
CREATE TABLE rijal_claims (
claim_id TEXT PRIMARY KEY,
page_slug TEXT NOT NULL,
claim_text TEXT NOT NULL,
narrator_chain JSONB NOT NULL,
chain_confidence NUMERIC(4,3),
valid_from TIMESTAMPTZ,
valid_until TIMESTAMPTZ,
superseded_by TEXT,
chain_status TEXT DEFAULT 'active'
);
CREATE TABLE narrator_registry (
narrator_id TEXT NOT NULL,
domain_tag TEXT NOT NULL,
narrator_type TEXT NOT NULL,
grade TEXT NOT NULL,
known_error_rate NUMERIC(4,3),
model_version TEXT,
is_active BOOLEAN DEFAULT TRUE,
PRIMARY KEY (narrator_id, domain_tag)
);
Because the pipeline's transformation points already emit trace identifiers, populating narrator_chain is an instrumentation task, not an architectural change.
Evaluation: What's Validated, What Failed, What's Measured#
ISNAD-Bench — Measured Against 1,200 Years of Ground Truth
The framework's newest and strongest validation. ISNAD-Bench grades 577,024 real hadith chains that classical scholars already graded by hand, then asks one falsifiable question: when ISNAD is given the scholars' own narrator grades, does its weakest-link rule reach the same chain verdicts they did?
The answer is reported as Cohen's κ — never raw accuracy, which is meaningless on a corpus dominated by ṣaḥīḥ.
| Comparison | Cohen's κ | Agreement |
|---|---|---|
| ISNAD vs scholarly consensus (strict default) | 0.8714 | 89.7% |
| ISNAD vs consensus (lenient opt-in) | 0.7610 | 82.4% |
| Human ceiling — a single scholar vs consensus | 0.450 | — |
| Human ceiling — scholar vs scholar | 0.331 | — |
| Shuffled-rank control | 0.047 | — |
| Corpus | 577,024 scholar-graded chains | |
Three things the benchmark taught
- Mutābaʿa was already load-bearing. The largest "disagreement" was chains scholars called weak-alone-but-ḥasan-if-corroborated. ISNAD said weak; they said acceptable. Checking the corpus, 88–92% of those chains have an independent corroborating route — the scholars were applying mutābaʿa, and it validated ISNAD's corroboration engine from the outside.
- Reliability has a timeline. 161 narrators declined over their lifetime (ikhtilāṭ), touching 35% of all chains. The scholars didn't discard them — they dated the decline. ISNAD's
get_grade_as_of()period-slicing, built on principle, turns out to be load-bearing across a third of the corpus. - The default was wrong. When a narrator had no grade (majhūl), ISNAD treated them as passable — generous, reasonable, wrong. κ = 0.761. The classical position caps unknown at ḍaʿīf. Flipping the default: κ = 0.871. Eleven points of κ from a 1,200-year-old method being right where modern intuition was wrong.
Full results, preregistered mapping, negative controls and the complete disagreement buckets: bench/docs/RESULTS.md · preregistered mapping.
RAGTruth — Hallucination Detection on Real LLM Output
A second, independent benchmark. On RAGTruth — 17,790 responses across 6 models — ISNAD's grounding critic catches 97.6% of hallucinations at κ = 0.575 (79.3% accuracy against a 71.4% baseline). Reported on ISNAD Cloud with the same preregistered discipline as ISNAD-Bench.
Setup
- 20,000 atomic claims extracted from real PDF text: OpenStax University Physics Vols 1–3 (CC BY 4.0) and Crowell's Light and Matter (CC BY-SA), split evenly, 10,000 per source.
- Domain tags: mechanics (7,600), general (6,890), electromagnetism (3,054), modern-quantum (1,728), optics-waves (728).
- Extraction via
deepseek-chatover chunked PDF text. - Split 30% calibration / 70% evaluation → 14,001 evaluation claims, across 10 random seeds.
- Four narrators with designed fault rates:
pdf-scraper@1.2(1%),ingest@good(2%),ingest@weak(15%),pdf-scraper@0.9-legacy(18%). - Faults: OCR noise, digit swap, sign flip, negation drop, unit corruption, formula mangling, entity swap, fabricated numerics, regime confusion. ~5% of chains marked incomplete.
- Review budgets: B ∈ {2%, 5%, 10%, 20%}.
Status at a Glance
| Component | Status | What remains |
|---|---|---|
| Weakest-link grading and quarantine | ✅ Validated | — |
| jarḥ–taʿdīl grade recovery | ⚠️ Partial | Recovered 3/4; missed the highest-fault narrator through insufficient per-cell calibration evidence |
| Corroboration (mutābaʿāt) | ✅ Validated | Negative controls on the physics corpus; organically weak chains |
| Confidence baseline | ⛔ Not tested | Baseline is synthetic; real model self-confidence untested |
| Matched-coverage (§8.4) | ✅ Measured | Chain-only 47.9% coverage @ 3.0% served-error; critic adds nothing on subtle corruption |
| Content criticism | ✅ Measured | LLM 1.000 recall / 0.000 false-consistent; offline tiers conservative |
| End-to-end serving pipeline | ⚠️ Partial | e2e utility measured (0.0% vs 50.0% served-error); live deployment open |
Validated: Weakest-Link Quarantine
Every claim whose chain contained a narrator graded rejected was quarantined — 4,057 claims, 29% of the evaluation split — and every quarantine was traceable to the specific narrator grade that caused it.
Partial Failure: The jarḥ–taʿdīl Loop
The loop recovered three of four narrator grades from audit evidence alone:
pdf-scraper@1.2(1% fault) → reliable in all 50 (narrator, domain) cellsingest@good(2%) → acceptable in all 50ingest@weak(15%) — deliberately left ungraded — independently discovered and driven to rejected in 49/50 cells, weak in one
pdf-scraper@0.9-legacy, carrying the highest designed fault rate (18%), never accumulated enough claims in the calibration split to earn any grade in any domain, and remained ungraded throughout. It was among the largest single sources of injected faults. The loop found the unreliable narrator it was given enough evidence about, and silently missed the one it was not.
Grade recovery is bounded by calibration coverage per (narrator, domain) cell, not only by the transition policy. Domain-conditioned grading multiplies the number of cells that must be filled, and a narrator rare within a domain is invisible to jarḥ–taʿdīl regardless of how unreliable it is. A deployment should monitor per-cell evidence counts as a first-class operational metric and treat a persistently under-evidenced cell as a risk, not a neutral absence.
Validated: Corroboration Across Three Corpora
| v1 (exact) | v2 (Wikipedia) | v3 (physics) | |
|---|---|---|---|
| Matching criterion | exact string | cosine ≥ 0.75 | cosine ≥ 0.80 |
| Corpus | 12 article intros | 30 article pairs | 2 textbooks |
| Sentences | 215 | 10,544 | 22,372 |
| Candidate pairs | 136 | 662 | 104 |
| Pairs evaluated | 136 | 603 | 104 |
| Match density | — | 6.3% | 0.5% |
| Corroboration fired | 68/136 | 603/603 | 104/104 |
| Negative controls | none run | 8/8 | none run |
| Independence basis | synthetic chains | separate editorial communities | different authors and publishers |
| Difficulty | easy | medium | hard |
The eight negative controls (v2) correctly produced no upgrade: no matching text; shared model family (the madār case); corroborators below the grade gate; a rejected base chain; the ṣaḥīḥ cap; shared upstream source; insufficient independent chains; empty corpus.
Matched-Coverage: The Measured Finding
The paper reported the matched-coverage comparison as "inconclusive" — ISNAD-gated coverage could not be driven above 4.8% with the reference critic. That question is now resolved as a measured, three-part finding (#126), not a shrug.
1. Chain grading + jarḥ–taʿdīl is the primary gate, and it works. With the critic treated as a no-op (acceptance riding entirely on the chain and narrator grades), ISNAD reaches 47.9% coverage at 3.0% served-error — the corrupt narrator is caught by the evidence loop, not by a content critic.
2. The matn critic is a rejection gate, not an acceptance gate. On a sample of the residual "leaks" (corrupted claims with a sound chain), the LLM critic caught 0 of 30 — subtle single-claim corruption is exactly what a content critic cannot reliably detect. Its false-consistent rate on meaning-changing corruption is 39.1% (#126), so it is unsafe as an acceptance gate.
3. The missing discriminator — corroboration — is structurally absent from this corpus. The classical defense against "sound chain, subtly wrong claim" is independent corroboration (mutābaʿāt). The §8 corpus has zero cross-source overlap, so corroboration fired 0 times by construction. Matched coverage on this corpus is structurally unattainable, not merely unfinished.
Full write-up: experiments/s8_gated_vs_ungated/results/MATCHED_COVERAGE_FINDING.md.
End-to-End Utility: The Fidelity→Utility Crossing
Every prior validation used synthetic fault injection or extracted sentences. Here the claims are generated by a live LLM (DeepSeek), labelled by a deterministic oracle (not the critic), then routed through the real pipeline.
| served-error | |
|---|---|
| ISNAD-gated | 0.0% |
| Ungated | 50.0% |
The gate served only the correct in-corpus restatements (with caveat) and routed all 12 novel-topic claims to review.
Full write-up: experiments/e2e_utility/RESULTS.md.
The Confidence Baseline Is Uninformative by Construction
A confidence-gated condition was included to represent current practice. It performs at chance — but for a reason that sharply limits what can be concluded from it. The corpus's model_confidence field is computed before fault injection and never updated. Confidence is statistically independent of corruption by construction. The point-biserial correlation is r = 0.04; mean confidence on corrupted claims (0.853) and clean claims (0.848) differ by less than one hundredth.
Honest Negatives: Transition Policy & Content Critic
Sweeping the downgrade threshold shows a direct tradeoff with no sweet spot:
| Downgrade threshold | Coverage | Served error | Review precision | Grades recovered |
|---|---|---|---|---|
| 3 | 7.1% | 0.0% | 0.11 | 2/4 |
| 6 | 9.4% | 0.0% | 0.15 | 2/4 |
| 10 | 10.0% | 0.0% | 0.11 | 1/4 |
| 15 | 10.0% | 0.0% | 0.13 | 0/4 |
| 25 | 10.0% | 0.0% | 0.13 | 0/4 |
There is no setting in this range that delivers both coverage and grade recovery. More audit evidence did not help — increasing the per-cell audit budget reduced coverage from 10.0% to 5.9% at 40 and above. The default policy over-penalises reliable narrators as evidence accumulates on small samples.
| Critic | Contra. recall | False-consistent |
|---|---|---|
LLMCritic (DeepSeek) | 1.000 | 0.000 |
LocalNLICritic | 0.760 | 0.000 |
HybridCritic | 0.720 | 0.040 |
EmbeddingCritic | 0.120 | 0.320 |
Only the LLM tier is near-perfect on genuine contradictions (100% recall, zero false-consistents); the offline tiers are safe but conservative. But on subtle meaning-changing corruption, the critic's false-consistent rate is 39.1% (#126) — it is a rejection gate, not an acceptance gate.
The earlier "~90% LLM / ~40% NLI / ~30% LocalNLI" figures were estimated on a template-injected word-swap set and have been retracted (issue #96); the measured hierarchy above replaces them.
The honest negative stands: with most real claims returning "unverifiable," the critic — not the grading mechanism — caps end-to-end coverage. The hierarchy is measured; the open piece is the hard case (a confident wrong claim within a covered domain).
Case Study: Contradiction Detection on Real Texts
A prototype compiled-wiki system ingested OpenStax University Physics Vols 1–3 and Crowell's Light and Matter via chunked PDF extraction, with DeepSeek-chat as the ingestion model, single run.
Result: 19 contradictions surfaced across the corpus, each manually reviewed and confirmed genuine — among them classical momentum (p = mv) versus photon momentum (p = h/λ); wave versus particle treatments of light; Newtonian versus relativistic kinematics.
All 19 were Type B — regime distinctions where both claims are true in their domains and resolution means qualification — not Type A errors. This validates the detection substrate and routing path; it does not demonstrate error-correction capability.
Limitations#
The analogy has edges.
Classical narrators were moral agents whose ʿadālah was a judgment of character formed over a lifetime; models are stochastic artifacts whose "integrity" is an engineering property of their deployment. The mapping is functional, not metaphysical. Classical grading also rested on scholarly consensus formed over generations; a registry populated by one team's evaluations is a far thinner epistemic base, and its grades should be presented with corresponding humility.
Registry cold start.
Weakest-link grading with an empty registry caps everything at ḥasan-tier or below. Arguably correct — new pipelines should earn trust — but it front-loads human review cost. Practical value depends on the jarḥ–taʿdīl loop converging faster than review budgets exhaust. §8 showed the intended mitigation (seed grading) goes less far than anticipated.
Grading humans.
Where human contributors are narrators, the registry becomes a reputation system over people, with the fairness, contestability, and workplace-power questions that implies. Flagged as requiring governance design — contestable grades, transparent criteria — not resolved.
False precision.
chain_confidence as a decimal invites over-trust. The ordinal-first principle is the mitigation; implementations that surface raw numbers to end users without tier context would be misusing the framework.
Independence is an idealisation.
Two chains may share no explicit narrator yet still fail together — if both terminate in the same answer model, or in models from a family with correlated training data. Classical scholars faced the same problem in the madār (the pivot narrator through which ostensibly independent chains converge). The reference implementation enforces an independence check on narrator identity, model family, and upstream source — but a genuinely correlated pair sharing none of those three would still pass, and content-level correlated error (two independent sources repeating the same received mistake) is not addressed at all.
Probabilistic and provisional claims.
The framework assumes claims with determinate truth values. "The Higgs boson mass is 125.1 GeV" carries an implicit uncertainty interval. Contradiction detection between a point estimate and an interval is not defined. The lifecycle columns provide a mechanism for supersession but not a semantics for uncertainty.
Storage and query overhead is estimated, not measured.
Chains are short (typically 3–5 links) so per-claim JSONB is small and bounded; the registry grows with pipeline changes, not corpus size. At the targeted scale — thousands to low millions of claims — neither table is expected to bottleneck. These are design expectations, not measured results. No latency or storage benchmark was run, and none is claimed.
Scope exclusions: Byzantine collusion among agents, cryptographic attestation of traces (complementary, not competing), and open-ended or creative generation. ISNAD grades claims that can in principle be true or false and checked against a corpus. It is a framework for knowledge accumulation, not for tasks where there is no fact of the matter to transmit.
What's Next#
Shipped Since the Paper's Submission Version
- ISNAD-Bench — measured against 577,024 real hadith chains. Cohen's κ = 0.8714 vs scholarly consensus (strict, 89.7% agreement), 0.7610 (lenient), shuffled control 0.047, human ceiling 0.331. Preregistered mapping + negative controls.
- Strict default — ungraded narrators (majhūl) now cap at ḍaʿīf, matching the classical position; leniency is opt-in via
lenient_unknown=True. - Multi-provider LLM critic — OpenRouter, OpenAI, DeepSeek, Anthropic, Gemini, Groq, Together, or local Ollama; measured 1.000 recall / 0.000 false-consistent on the committed eval set.
- Audit evidence layer — tamper-evident
AuditRecordwith RFC 8785-canonical SHA-256, DAG support, linear hash chain + Merkle batch log, detached signatures (HMAC + Ed25519), PII redaction, and governance mapping. - Period-sliced grades (
get_grade_as_of()) — the ikhtilāṭ remedy, validated by ISNAD-Bench M4. - Per-role precision (#3) — precision (ḍabṭ) graded per (narrator, role, domain); integrity stays per-narrator. Integrity ladder: ʿadālah never forgets, ḍabṭ recovers.
- Evidence-backed seed grades —
Registry.seed()/seed_from_benchmark()record evidence, not bare grades; seeds survive the Bayesian posterior. - Chain independence (#125) — document-hash madār detection and a disjointness discount; corroboration is weighted by independence score, and provenance is surfaced in results/API.
- Working content critics:
EmbeddingCritic,HybridCritic(NLI),LLMCritic - LangChain integration:
IsnadTracer+seed_registry+IsnadMiddleware(gates claims); plus LangGraph, CrewAI, LlamaIndex, OpenTelemetry ingest. - MCP integration — grade MCP servers as narrators:
MCPToolObserverrecords tool calls as TOOL links, and agrade_claimtool exposes the operator's registry to agents (duck-typed; no MCP SDK required). - JS/TS audit-record verifier on npm —
npm i isnadverifies SHA-256 / Merkle / detached-signature integrity of audit records (verifier only; grading stays in Python). - FastAPI service, CLI, SQLAlchemy persistence, endpoint/version-drift handling
- Live Verify integration — consume a cryptographic seal as a high-trust narrator
- Model-drift leaderboard (#71) — a preregistered, LLM-free harness measuring hallucination across chain depth 1–5, with a deterministic offline drift injector and negative controls. First live result: hallucination originates at memory-generation (hop 1) and stays flat across depth — the relay preserves the error, it doesn't compound it.
- Chain-scoped content grounding (#216) —
chain_scoped_corpus()+ a swappableGroundingPolicyflag a claim grounded only off its own chain (a grounding gap, not proven contamination). - Content-madār tightening (#232) — shared-error fingerprinting now requires error-distinctive signals (number+unit, citation+number/date); false-positive-on-agreement down from 0.750 to 0.375 with token-bearing recall held at 1.0.
- Warm default registry +
isnad scan(#203/#204/#206) — an evidence-sourced seed set (~39 narrators) so a fresh pipeline no longer starts fully ungraded, plus a CLI that maps live transmitters onto the registry (vouched vs cold).
Highest-Priority Open Work (collaborators welcome)
- The hard case. ISNAD-Bench validates the deterministic grading rule, and the end-to-end utility experiment (#128) shows the gated pipeline serving 0.0% vs 50.0% served-error on live LLM-generated claims. What remains is the hard case: a confident wrong claim within a covered domain, where "review" is not the easy call.
- A live production deployment where chains are written automatically at ingest and the registry gates a serving path under real query load.
- Chain independence — the hardest open problem: topology cannot prove two chains are independent. Tracking issue →
- Negative controls on the physics corroboration corpus (v3) and corroboration tested on organically weak chains.
- Per-cell evidence-count monitoring as a first-class operational metric.
- Extension to probabilistic and provisional claims.
Good First Issues
- An alternative critic (sentence-transformers embedding, CrewAI integration)
- A seed-grade bootstrapper from published benchmark accuracies
- A new
CorrelationDetectorusing embedding similarity or model-card lineage - Extend semantic corroboration to multi-source corpora
- A calibrated
TransitionPolicyfrom your own pipeline's §8 experiment data - Domain-specific
ContentCriticwith formula canonicalisation - Pipeline adapters for CrewAI or AutoGen tracing
In the Wild#
Live Verify — Independent Comparison
Paul Hammant — co-creator of Selenium (~22 years ago), maintainer of trunkbaseddevelopment.com, ex-ThoughtWorks, based in Scotland. Emailed Ali directly after finding the repo: "i saw your isnad repo, and looks a very interesting approach.. i'd love contributing to it incase i have the time."
A dedicated comparison document: Live Verify vs. Isnād–Rijāl
Both attack the same root problem — why should I trust this claim that reached me through a chain of hands? — and both reach for a provenance answer rather than a plausibility answer. But they sit at opposite ends of the trust pipeline. They are complementary rather than competing.
| ISNAD | Live Verify | |
|---|---|---|
| Layer | Inside an AI pipeline, per-claim | Between organisations, per-artifact |
| Mechanism | Graded reputation registry + Bayesian loop | Cryptographic hash + domain endpoint |
| Trust root | Accumulated narrator track record | An authorizedBy chain to a jurisdiction's root namespace |
| Output | serve / review / quarantine | verified / not-verified + status |
| Truth stance | Probabilistic (how reliable) | Binary (unaltered vs. altered) |
| Handles | Distortion by AI transformers | Forgery/tampering of documents |
"ISNAD's README carries a 'What's Validated vs. What's Not' box and reports two of its own results as inconclusive rather than positive. … Both projects treat scoping-out over-claims as a first-class feature, not an afterthought." — from the Live Verify comparison document.
Read the full comparison → · Have your own project or comparison? Add it — reach out.
Listed in LangChain's Middleware Directory
ISNAD ships as a community middleware integration in LangChain's official docs. In the "All middleware" directory — sorted by monthly downloads — it ranks sixth, right after OpenAI, Anthropic, AWS, Microsoft Foundry, and CopilotKit.
Indexing, Archiving & Aggregation
Also discoverable via NASA ADS, Google Scholar, Semantic Scholar, Connected Papers, alphaXiv, CatalyzeX, Litmaps, scite, OpenAIRE, and Hugging Face Daily Papers. Reddit launch threads — URLs pending; will be linked when recovered.
Prior & Adjacent Work
- PROV-AGENT (Souza et al., IEEE e-Science 2025) — extends W3C PROV to capture agent interactions.
- Agent-Sentry (Sequeira et al., 2026) — bounding LLM agents via execution provenance.
- HadithRank (Hawramani) — models isnād topology using information theory, treating independent chains as redundant channels. That a mathematician grading hadith and this paper grading agents converge on the same redundancy rule is evidence the mapping is structural, not decorative.
- The Provenance Paradox (Prakash, 2026) — when delegates can inflate self-reported quality, quality-based routing can perform worse than random. Consistent with ISNAD's design intuition but not a replication.
- RAPS (2026) — reputation-aware publish–subscribe with Bayesian overlays.
- TrustTrade (2026) — replaces "uniform trust" with cross-agent consistency weighting.
- AgentPoison (Chen et al., NeurIPS 2024) — red-teaming LLM agents by poisoning memory/knowledge bases; the demonstrated threat the mawḍūʿ tier defends against.
- Karpathy's "LLM Wiki" sketch (2026) — the compiled-knowledge-base system class that motivates the whole framework.
Two public repositories independently reference isnad-style provenance chains: cognalith/isnad (agent skill security with provenance verification) and VAMaksimov/threat-intel-format (threat intelligence schema with isnad-style provenance chains). Neither is affiliated.
Links#
Cite#
Paper
@article{raja2026grading,
author = {Ali Zahid Raja},
title = {Grading the Narrators: An Isnad--Rijal Framework for
Claim-Level Provenance in Multi-Agent Knowledge Systems},
year = 2026,
doi = {10.48550/arXiv.2607.24117},
eprint = {2607.24117},
archivePrefix = {arXiv},
primaryClass = {cs.AI}
}
Software
@software{raja2026isnad,
author = {Ali Zahid Raja},
title = {Isnad--Rijal Framework: Reference Implementation},
year = 2026,
doi = {10.5281/zenodo.21216872},
orcid = {0009-0003-7875-4590}
}
LLM usage disclosure: drafting and editing were assisted by a large language model; all technical claims, design decisions, experimental results, and final text were reviewed by and are the responsibility of the author.
Frequently Asked Questions#
What is ISNAD in one sentence?
ISNAD is an open-source framework that grades every agent, scraper, and model in a claim's transmission chain — adapted from classical Islamic hadith science — to tell you how much to trust a claim that has passed through multiple AI hands.
How is this different from W3C PROV or PROV-AGENT?
PROV records what happened — it captures chains, agents, and activities. ISNAD grades who did it and whether to trust the result. PROV gives you a transmission log; ISNAD tells you which link in that log is the weakest, and what action to take. See the positioning tables for a detailed comparison.
Does ISNAD outperform confidence-based gating?
The matched-coverage question is now a measured finding, not an open one. Chain grading alone reaches 47.9% coverage at 3.0% served-error (no content critic); the matn critic adds nothing on subtle single-claim corruption (0/30 caught); and corroboration — the one mechanism that could have discriminated — is structurally absent from the §8 corpus. See the measured finding.
What does "the chain is graded by its weakest link" mean in practice?
A ṣaḥīḥ model at the end of a chain cannot rescue a ḍaʿīf scraper at the start. The chain's grade is the minimum grade among its narrators — adjusted by transformation type (destructive vs. generative) — because information lost early in the chain cannot be recovered downstream. The chain animation demonstrates this visually.
What is mutābaʿāt / corroboration, and when does it not fire?
Mutābaʿāt is the mechanism by which independent chains asserting the same claim corroborate each other, upgrading the claim's grade (capped at ṣaḥīḥ). It does not fire when chains share a narrator, model family, or upstream source — the classical madār case. It is validated at 603/603 on Wikipedia and 104/104 on physics textbooks, with 8/8 negative controls passing. No negative controls have been run on the physics corpus yet.
Why grade narrators per domain instead of globally?
A model precise on historical dates may be unreliable on quantum-mechanical equations. The registry key is (narrator, domain), not narrator alone. This mirrors classical practice — rijāl scholars graded transmitters domain-specifically as a matter of course.
What happens when a model version changes?
A model-version bump resets the narrator to ungraded until re-evaluated. Version drift is treated as a new narrator — track record does not carry forward. This is intentional: a model at v2 may behave differently from the same model at v1, and inherited reputation would be misleading. Known blind spot: hosted models whose weights change without a version-string change will silently retain a grade they no longer merit.
What is the "unverifiable" verdict and why does it matter?
"Unverifiable" is the third column of the decision matrix — the classical position of tawaqquf (suspension of judgment). A content critic that cannot evaluate a claim must be able to say so rather than defaulting to "consistent." Routing unverifiable claims conservatively is what makes the framework safe under a weak critic. It is also the binding constraint on practical coverage — most real claims receive this verdict under the reference critic.
What did the evaluation fail to show?
Three things, reported in the same typographic weight as the successes: (1) The jarḥ–taʿdīl loop recovered 3 of 4 narrator grades but missed the highest-fault narrator through insufficient calibration evidence. (2) The confidence baseline is synthetic and proves nothing about real LLM self-confidence. (3) The matn critic's false-consistent rate on meaning-changing corruption is 39.1% (#126) — it is a rejection gate, not an acceptance gate. See the evaluation section.
Is this a religious claim or ruling?
No. ISNAD borrows a methodology — the epistemological architecture of isnād criticism — from classical Islamic hadith science. It makes no religious claims, issues no rulings, and does not position any computational system as an authority within the religious sciences. See the positionality statement.
Can I use this in production today?
The offline pipeline has been exercised end to end on 20,000 claims, and the end-to-end utility experiment (#128) runs live LLM-generated claims through the real pipeline — gated serving at 0.0% vs 50.0% ungated served-error. The LangChain integration is shipped and tested. A live production deployment where chains are written automatically at ingest and the registry gates a serving path under real query load is still open work.
How do I cite it?
Two BibTeX entries are available in the Cite section — one for the arXiv paper and one for the software. Paper: arXiv:2607.24117 (cs.AI). Software DOI: 10.5281/zenodo.21216872.
What licence is it under?
Code: Apache-2.0 — with a public commitment that src/isnad/ stays Apache-2.0 in perpetuity (see LICENSING.md). Paper and documentation: CC BY 4.0.
How can I contribute?
The CONTRIBUTING.md lists good first issues — alternative critics, seed-grade bootstrappers, correlation detectors, and pipeline adapters. The highest-impact open problem is chain independence (#54). Pull requests welcome.