ISNAD
Claim-level trust for multi-agent AI systems.
Every claim carries its chain. Every transmitter is graded. Adapted from 1,200 years of hadith transmission science.
The Living Chain#
Every claim carries its chain.
Try the Decision Matrix#
Select a chain grade and content verdict — the matrix routes your claim.
Show full 4×3 decision matrix
| Consistent | Contradiction | Unverifiable | |
|---|---|---|---|
| Ṣaḥīḥ-tier | Serve directly; cache | ʿIlal flag → human review | Serve with caveat |
| Ḥasan-tier | Serve with caveat; seek corroboration | Hold; do not serve | Hold in review queue |
| Ḍaʿīf-tier | Hold; seek corroboration | Quarantine | Hold in review queue |
| Mawḍūʿ-tier | Reject & quarantine narrator | Reject & quarantine narrator | Reject & quarantine narrator |
Corroboration & the Madār#
Two chains, same claim. Does corroboration fire — or does the madār block it?
The Problem#
When a multi-agent AI system answers you, that claim has passed through many hands — a source, a scraper, an ingestion model, a synthesis model. Each hand can drop, distort, or invent.
A single factual claim may pass through document retrieval, extraction, summarization, knowledge integration, and answer synthesis before reaching a user. Provenance systems can reconstruct that chain after the fact. They provide limited guidance on the question users actually care about: this specific claim, having passed through these specific transformers — should I trust it?
Provenance tools record what happened. Nothing grades who did it — or tells you how much to trust the result. A confident model at the end of a chain cannot repair a corrupted extraction at the start of it.
The Idea#
Classical Islamic hadith science spent twelve centuries on a structurally identical problem: whether to trust knowledge transmitted through chains of human narrators. ISNAD transfers that methodology to AI pipelines.
Every claim carries its transmission chain (isnād)
Every transmitter keeps a living, per-domain grade (rijāl)
The weakest link caps the chain's trust
Independent corroboration upgrades — capped and gated (mutābaʿāt)
Content is criticized separately from the chain (matn)
Concept Mapping: From Ḥadīth Science to AI Pipelines#
The intellectual heart of the framework — the paper's Table 1, mapping each classical term to its ISNAD usage.
| Term | Classical meaning | Usage in ISNAD |
|---|---|---|
| isnād | The complete chain of transmitters attached to a narration | The agent chain attached to a claim |
| matn | The transmitted content itself | The claim text |
| rijāl | "The transmitters"; the discipline of evaluating them | The narrator registry |
| ʿadālah | A narrator's integrity / uprightness | Resistance to manipulation; source trustworthiness |
| ḍabṭ | A narrator's precision and retention accuracy | Measured error rate of an agent/model |
| al-jarḥ wa-l-taʿdīl | Formal criticism and accreditation of narrators | The registry update process (evals, audits) |
| ittiṣāl / munqaṭiʿ | Chain continuity / a broken chain | Complete trace / missing trace segment |
| ṣaḥīḥ, ḥasan, ḍaʿīf, mawḍūʿ | Sound, good, weak, fabricated (grades) | Trust tiers for claims |
| mutābaʿāt | Independent corroborating chains | Multiple disjoint agent chains asserting one claim |
| ʿilal | Hidden defects: sound-looking chain, defective content | Chain passes but content criticism fails |
| muḥaddith | The expert scholar who adjudicates difficult cases | The human reviewer in the loop |
Refining the Weakest-Link Rule
In hadith transmission a weak link is unconditionally corrupting. AI transformations are not all alike. A destructive transformation — extraction, chunking, lossy summarization — can only lose or corrupt information, and there the strict minimum applies: nothing downstream recovers what the extractor dropped. A generative transformation can either repair upstream noise or introduce fresh corruption by favouring its parametric prior over retrieved evidence. For generative steps the rule is therefore not a blind minimum but a bounded one: a generative narrator can raise the floor only up to its own grade, and only when corroboration supports the repair — and it can always lower it.
This mirrors a distinction the muḥaddithūn drew themselves, between riwāya bi-l-lafẓ (transmission of exact wording) and riwāya bi-l-maʿnā (transmission by meaning). The weakest-link principle survives; it is refined by transformation type.
(narrator, domain), not narrator alone. A model precise on historical dates may be unreliable on quantum-mechanical equations. This is a return to classical practice — rijāl scholars graded transmitters domain-specifically as a matter of course.
The Decision Matrix#
The framework's most actionable artifact — chain grade × content verdict → action. The third column is the binding constraint on real-world coverage.
| Matn: consistent | Matn: contradiction | Matn: unverifiable | |
|---|---|---|---|
| Ṣaḥīḥ-tier chain complete; all narrators graded reliable | Serve directly; cache | ʿIlal flag → human review. Highest-value review case. | Serve with caveat |
| Ḥasan-tier chain complete; ≥1 narrator ungraded or mid-tier | Serve with explicit confidence caveat; seek corroboration | Hold in review queue; do not serve | Hold in review queue |
| Ḍaʿīf-tier chain weak narrator, or munqaṭiʿ | Hold; seek corroborating chain before serving | Quarantine | Hold in review queue |
| Mawḍūʿ-tier chain rejected narrator — e.g. known injection source | Reject and log; quarantine the narrator | Reject and log; quarantine the narrator | Reject and log; quarantine the narrator |
Contradictions go to humans by default. Knowledge-conflict research consistently finds LLMs unreliable at representing and reconciling competing evidence, and during prototype development LLM synthesis handled contradiction cases worse than an extractive baseline until contradiction markers were made explicitly visible in context — after which auto-resolution was still left disabled in favour of gated review.
Worked Example: End to End#
Claim: "the momentum of a photon is p = h/λ", ingested from an openly licensed physics text.
Chain: [OpenStax University Physics Vol. 3 — source; publisher-trusted, ʿadālah high] → [pdf-scraper v1.2 — ḍabṭ graded: high extraction fidelity on text-layer PDFs] → [ingest-analysis model M@v — ungraded, recent version bump] → [ingest-renderer M@v — ungraded]
Chain grade: complete (ittiṣāl holds), but two narrators ungraded → ḥasan-tier by the weakest-link rule.
Matn criticism: the corpus already contains a page asserting momentum as p = mv without regime qualification. Contradiction detected.
Matrix cell: ḥasan × contradiction → review queue; do not serve.
Adjudication: a human reviewer recognises both claims as valid in different regimes — classical and quantum — and updates both pages with regime qualifiers, resolving the contradiction once, at ingest time, rather than leaving every future query to renegotiate it. The resolution feeds back: the ingestion model's contradiction-detection behaviour on this case becomes taʿdīl evidence toward its grade.
This example is implemented as a runnable demo and an integration test in the reference implementation (make demo, examples/worked_example.py, tests/test_worked_example.py).
What ISNAD Adds Over Existing Provenance#
Each Mechanism Against Its Strongest Existing Ancestor
| Mechanism | Strongest ancestor | What ISNAD adds |
|---|---|---|
| Chain aggregation | Jøsang subjective logic (2001) — trust discounting along chains | Transformation typing (riwāya bi-l-lafẓ vs. bi-l-maʿnā); completeness semantics |
| Source registry | Truth discovery (Yin et al. 2008; Dong et al. 2014) | Per-domain grading; version-bump resets; quarantine lifecycle |
| Content criticism | FEVER-style fact-checking | Decoupled from chain quality; operates on matn independently; explicit unverifiable verdict |
| Agent grading | MAS reputation (EigenTrust, ReGreT, FIRE) | Claim-level provenance, not just partner selection |
| Decision routing | — (no direct ancestor) | Chain grade × content verdict → concrete serve/review/quarantine action |
Capability Coverage
| Approach | Chain capture | Graded transmitters | Content criticism | Decision routing |
|---|---|---|---|---|
| W3C PROV / PROV-AGENT | ✅ | ❌ | ❌ | ❌ |
| Truth discovery (TruthFinder, KBT) | ❌ | ✅ | partial | ❌ |
| MAS reputation (EigenTrust, ReGreT) | ❌ | ✅ | ❌ | partial |
| FEVER-style fact-checking | ❌ | ❌ | ✅ | ❌ |
| ISNAD | ✅ | ✅ | ✅ | ✅ |
Chain capture means claim-level transmission chains are recorded per claim; graded transmitters means each transformer carries a maintained reliability grade; content criticism means content is evaluated independently of chain quality; decision routing means the system emits a concrete serve/review/quarantine action. A ✅ indicates the approach addresses that capability as a first-class concern.
Deep lineage: Condorcet's jury theorem (1785), Hume on testimony (1748), the hearsay-within-hearsay rule (US FRE 805), Dempster–Shafer evidence combination, Dawid & Skene (1979), Jøsang's subjective logic (2001), TruthFinder (2008), Knowledge Vault (2014), Knowledge-Based Trust (2015). The paper's claim is not novelty for its own sake but the transfer of the most operationally complete solution to a domain that lacks it.
Quickstart#
Install
pip install isnad
60-Second Example
from isnad import Registry, Chain, ChainLinkSpec, grade_chain, decide
from isnad.types import NarratorGrade, ContentVerdict
from isnad.critics import EmbeddingCritic
# Build a chain: source → scraper → model
chain = Chain([
ChainLinkSpec("openstax-textbook", 0, domain="physics"),
ChainLinkSpec("pdf-scraper-v2", 1),
ChainLinkSpec("ingest-model-v3", 2),
])
# Seed-grade known narrators
reg = Registry()
reg.register("openstax-textbook", "physics", grade=NarratorGrade.RELIABLE)
reg.register("pdf-scraper-v2", "physics", grade=NarratorGrade.RELIABLE)
reg.register("ingest-model-v3", "physics", grade=NarratorGrade.ACCEPTABLE)
# Grade the chain
grades = [reg.get_grade(l.narrator_id, l.domain) for l in chain.links]
transforms = [l.transform_type for l in chain.links]
chain_grade = grade_chain(grades, transforms, is_complete=True)
# Content criticism — decoupled from chain grading
critic = EmbeddingCritic()
verdict = critic.evaluate("p = h/λ", "p = h/lambda", ["p = mv"])
action = decide(chain_grade, verdict)
print(f"Chain: {chain_grade.value.upper()} | Content: {verdict.value} | Action: {action.value}")
LangChain Integration (5 Lines)
pip install isnad[langchain]
from isnad.integrations.langchain import IsnadTracer, seed_registry
from isnad.critics import EmbeddingCritic
reg = seed_registry({"source:docs": "reliable", "model:gpt-4o": "acceptable"})
tracer = IsnadTracer(registry=reg, critic=EmbeddingCritic())
chain.invoke("What is F=ma?", config={"callbacks": [tracer]})
print(tracer.report())
Run the Experiments
git clone https://github.com/alizahidraja/isnad && cd isnad make install # uv sync make test # zero config, SQLite fallback make demo # the paper's worked example (§4.5) make check # lint + type-check + test
No database required for pure-logic tests; PostgreSQL is optional (docker compose up, set ISNAD_DATABASE_URL).
Architecture & Module Map#
| Concept | What it does | Module |
|---|---|---|
| isnād (chain) | Ordered, gap-checked transmission chain per claim | isnad/core/chain.py |
| rijāl (registry) | Graded narrator store per (alias@version, domain) | isnad/core/registry.py |
| jarḥ–taʿdīl | Evidence-driven state machine for narrator grades | isnad/core/registry.py |
| Bayesian grading | Beta-distribution narrator grades (default) | isnad/core/registry.py |
| ittiṣāl | Completeness as epistemic property (gap → ḍaʿīf) | isnad/core/chain.py |
| Weakest-link grading | Chain grade = refined minimum over narrators | isnad/core/grading.py |
| mutābaʿāt (corroboration) | Independent-chain upgrade + madār detection | isnad/core/corroboration.py |
| matn criticism | Content evaluated independently of chain quality | isnad/critics/ |
| Decision matrix | Chain × content → action router | isnad/core/decision.py |
| Persistence | SQLAlchemy-backed registry (swap via protocol) | isnad/storage/ |
| API | FastAPI service with DI + Prometheus metrics | isnad/api/ |
| CLI | isnad serve, isnad seed | isnad/cli/ |
| ʿadālah / ḍabṭ | Integrity and precision as two distinct axes | isnad/types.py |
Pluggable Strategies — Updated Defaults
| Strategy | Protocol | Default | What it decides |
|---|---|---|---|
GradingStrategy | isnad/types.py | RefinedWeakestLink | How link grades combine into a chain grade |
TransitionPolicy | isnad/types.py | BayesianTransitionPolicy | How evidence moves narrators between states |
CorroborationPolicy | isnad/types.py | CappedCorroborationPolicy | How independent chains upgrade a claim |
CorrelationDetector | isnad/types.py | SharedLineageDetector | Whether two chains are truly independent |
ContentCritic | isnad/critics/base.py | HybridCritic / EmbeddingCritic | Content contradiction detection |
Reference Schema
CREATE TABLE rijal_claims (
claim_id TEXT PRIMARY KEY,
page_slug TEXT NOT NULL,
claim_text TEXT NOT NULL,
narrator_chain JSONB NOT NULL,
chain_confidence NUMERIC(4,3),
valid_from TIMESTAMPTZ,
valid_until TIMESTAMPTZ,
superseded_by TEXT,
chain_status TEXT DEFAULT 'active'
);
CREATE TABLE narrator_registry (
narrator_id TEXT NOT NULL,
domain_tag TEXT NOT NULL,
narrator_type TEXT NOT NULL,
grade TEXT NOT NULL,
known_error_rate NUMERIC(4,3),
model_version TEXT,
is_active BOOLEAN DEFAULT TRUE,
PRIMARY KEY (narrator_id, domain_tag)
);
Because the pipeline's transformation points already emit trace identifiers, populating narrator_chain is an instrumentation task, not an architectural change.
Evaluation: What's Validated, What Failed, What's Inconclusive#
Setup
- 20,000 atomic claims extracted from real PDF text: OpenStax University Physics Vols 1–3 (CC BY 4.0) and Crowell's Light and Matter (CC BY-SA), split evenly, 10,000 per source.
- Domain tags: mechanics (7,600), general (6,890), electromagnetism (3,054), modern-quantum (1,728), optics-waves (728).
- Extraction via
deepseek-chatover chunked PDF text. - Split 30% calibration / 70% evaluation → 14,001 evaluation claims, across 10 random seeds.
- Four narrators with designed fault rates:
pdf-scraper@1.2(1%),ingest@good(2%),ingest@weak(15%),pdf-scraper@0.9-legacy(18%). - Faults: OCR noise, digit swap, sign flip, negation drop, unit corruption, formula mangling, entity swap, fabricated numerics, regime confusion. ~5% of chains marked incomplete.
- Review budgets: B ∈ {2%, 5%, 10%, 20%}.
Status at a Glance
| Component | Status | What remains |
|---|---|---|
| Weakest-link grading and quarantine | ✅ Validated | — |
| jarḥ–taʿdīl grade recovery | ⚠️ Partial | Recovered 3/4; missed the highest-fault narrator through insufficient per-cell calibration evidence |
| Corroboration (mutābaʿāt) | ✅ Validated | Negative controls on the physics corpus; organically weak chains |
| Confidence baseline | ⛔ Not tested | Baseline is synthetic; real model self-confidence untested |
| Matched-coverage advantage | ❓ Inconclusive | Requires a critic that can reach matched coverage |
| Content criticism | ⚠️ Partial | Semantic or LLM-backed critic |
| End-to-end serving pipeline | 🔲 Open | Live deployment at practical coverage |
Validated: Weakest-Link Quarantine
Every claim whose chain contained a narrator graded rejected was quarantined — 4,057 claims, 29% of the evaluation split — and every quarantine was traceable to the specific narrator grade that caused it.
Partial Failure: The jarḥ–taʿdīl Loop
The loop recovered three of four narrator grades from audit evidence alone:
pdf-scraper@1.2(1% fault) → reliable in all 50 (narrator, domain) cellsingest@good(2%) → acceptable in all 50ingest@weak(15%) — deliberately left ungraded — independently discovered and driven to rejected in 49/50 cells, weak in one
pdf-scraper@0.9-legacy, carrying the highest designed fault rate (18%), never accumulated enough claims in the calibration split to earn any grade in any domain, and remained ungraded throughout. It was among the largest single sources of injected faults. The loop found the unreliable narrator it was given enough evidence about, and silently missed the one it was not.
Grade recovery is bounded by calibration coverage per (narrator, domain) cell, not only by the transition policy. Domain-conditioned grading multiplies the number of cells that must be filled, and a narrator rare within a domain is invisible to jarḥ–taʿdīl regardless of how unreliable it is. A deployment should monitor per-cell evidence counts as a first-class operational metric and treat a persistently under-evidenced cell as a risk, not a neutral absence.
Validated: Corroboration Across Three Corpora
| v1 (exact) | v2 (Wikipedia) | v3 (physics) | |
|---|---|---|---|
| Matching criterion | exact string | cosine ≥ 0.75 | cosine ≥ 0.80 |
| Corpus | 12 article intros | 30 article pairs | 2 textbooks |
| Sentences | 215 | 10,544 | 22,372 |
| Candidate pairs | 136 | 662 | 104 |
| Pairs evaluated | 136 | 603 | 104 |
| Match density | — | 6.3% | 0.5% |
| Corroboration fired | 68/136 | 603/603 | 104/104 |
| Negative controls | none run | 8/8 | none run |
| Independence basis | synthetic chains | separate editorial communities | different authors and publishers |
| Difficulty | easy | medium | hard |
The eight negative controls (v2) correctly produced no upgrade: no matching text; shared model family (the madār case); corroborators below the grade gate; a rejected base chain; the ṣaḥīḥ cap; shared upstream source; insufficient independent chains; empty corpus.
Inconclusive: The Matched-Coverage Comparison
This is the result most projects would hide. Lead with it.
Comparing served-error rates across conditions operating at different coverage levels is not meaningful — a system serving 5% of claims will trivially show lower error than one serving 100%. The standard remedy is a matched-coverage sweep. It was specified in advance, attempted, and could not be completed.
With the reference content critic, ISNAD-gated serving could not be driven above 4.8% coverage at any review budget between 5% and 50%. The risk–coverage relationship is a single point, not a curve.
| Target coverage | ISNAD coverage | ISNAD error | Baseline coverage | Baseline error |
|---|---|---|---|---|
| 20% | 4.8% (not reached) | 0.0% | 20.0% | 15.1% |
| 30% | 4.8% (not reached) | 0.0% | 30.0% | 15.8% |
| 50% | 4.8% (not reached) | 0.0% | 50.0% | 16.3% |
| 70% | 4.8% (not reached) | 0.0% | 70.0% | 16.0% |
| 90% | 4.8% (not reached) | 0.0% | 90.0% | 16.1% |
The rows are reported to show what was attempted and where it failed, not as an advantage.
The Confidence Baseline Is Uninformative by Construction
A confidence-gated condition was included to represent current practice. It performs at chance — but for a reason that sharply limits what can be concluded from it. The corpus's model_confidence field is computed before fault injection and never updated. Confidence is statistically independent of corruption by construction. The point-biserial correlation is r = 0.04; mean confidence on corrupted claims (0.853) and clean claims (0.848) differ by less than one hundredth.
Honest Negatives: Transition Policy & Content Critic
Sweeping the downgrade threshold shows a direct tradeoff with no sweet spot:
| Downgrade threshold | Coverage | Served error | Review precision | Grades recovered |
|---|---|---|---|---|
| 3 | 7.1% | 0.0% | 0.11 | 2/4 |
| 6 | 9.4% | 0.0% | 0.15 | 2/4 |
| 10 | 10.0% | 0.0% | 0.11 | 1/4 |
| 15 | 10.0% | 0.0% | 0.13 | 0/4 |
| 25 | 10.0% | 0.0% | 0.13 | 0/4 |
There is no setting in this range that delivers both coverage and grade recovery. More audit evidence did not help — increasing the per-cell audit budget reduced coverage from 10.0% to 5.9% at 40 and above. The default policy over-penalises reliable narrators as evidence accumulates on small samples.
A separate evaluation of the reference critics on 20 labelled claims returned perfect precision and recall — but those contradictions were template-injected and built from patterns the critics already encode. The numbers overstate real capability and are not reported as evidence.
A semantic or LLM-backed critic is the single highest-value component for practical end-to-end coverage, and building one is the most immediate item of future work.
Case Study: Contradiction Detection on Real Texts
A prototype compiled-wiki system ingested OpenStax University Physics Vols 1–3 and Crowell's Light and Matter via chunked PDF extraction, with DeepSeek-chat as the ingestion model, single run.
Result: 19 contradictions surfaced across the corpus, each manually reviewed and confirmed genuine — among them classical momentum (p = mv) versus photon momentum (p = h/λ); wave versus particle treatments of light; Newtonian versus relativistic kinematics.
All 19 were Type B — regime distinctions where both claims are true in their domains and resolution means qualification — not Type A errors. This validates the detection substrate and routing path; it does not demonstrate error-correction capability.
Limitations#
The analogy has edges.
Classical narrators were moral agents whose ʿadālah was a judgment of character formed over a lifetime; models are stochastic artifacts whose "integrity" is an engineering property of their deployment. The mapping is functional, not metaphysical. Classical grading also rested on scholarly consensus formed over generations; a registry populated by one team's evaluations is a far thinner epistemic base, and its grades should be presented with corresponding humility.
Registry cold start.
Weakest-link grading with an empty registry caps everything at ḥasan-tier or below. Arguably correct — new pipelines should earn trust — but it front-loads human review cost. Practical value depends on the jarḥ–taʿdīl loop converging faster than review budgets exhaust. §8 showed the intended mitigation (seed grading) goes less far than anticipated.
Grading humans.
Where human contributors are narrators, the registry becomes a reputation system over people, with the fairness, contestability, and workplace-power questions that implies. Flagged as requiring governance design — contestable grades, transparent criteria — not resolved.
False precision.
chain_confidence as a decimal invites over-trust. The ordinal-first principle is the mitigation; implementations that surface raw numbers to end users without tier context would be misusing the framework.
Independence is an idealisation.
Two chains may share no explicit narrator yet still fail together — if both terminate in the same answer model, or in models from a family with correlated training data. Classical scholars faced the same problem in the madār (the pivot narrator through which ostensibly independent chains converge). The reference implementation enforces an independence check on narrator identity, model family, and upstream source — but a genuinely correlated pair sharing none of those three would still pass, and content-level correlated error (two independent sources repeating the same received mistake) is not addressed at all.
Probabilistic and provisional claims.
The framework assumes claims with determinate truth values. "The Higgs boson mass is 125.1 GeV" carries an implicit uncertainty interval. Contradiction detection between a point estimate and an interval is not defined. The lifecycle columns provide a mechanism for supersession but not a semantics for uncertainty.
Storage and query overhead is estimated, not measured.
Chains are short (typically 3–5 links) so per-claim JSONB is small and bounded; the registry grows with pipeline changes, not corpus size. At the targeted scale — thousands to low millions of claims — neither table is expected to bottleneck. These are design expectations, not measured results. No latency or storage benchmark was run, and none is claimed.
Scope exclusions: Byzantine collusion among agents, cryptographic attestation of traces (complementary, not competing), and open-ended or creative generation. ISNAD grades claims that can in principle be true or false and checked against a corpus. It is a framework for knowledge accumulation, not for tasks where there is no fact of the matter to transmit.
What's Next#
Shipped Since the Paper's Submission Version
- Bayesian (Beta-distribution) narrator grading as the default transition policy, replacing hardcoded thresholds;
ISNAD_POLICYenv override - Working content critics:
EmbeddingCritic(TF-IDF, offline),HybridCritic(NLI),LLMCritic - Seed-grade bootstrapping — validated as improving coverage from ~5% to ~10%;
ISNAD_SEED_CONFIGenv var - LangChain integration:
IsnadTracercallback handler +seed_registry, 9 integration tests passing - FastAPI service with dependency injection and Prometheus metrics; CLI
- SQLAlchemy-backed persistence behind a swappable storage protocol
- Endpoint identity / model-version-drift handling (
alias@versionkeying) - Corroboration v3 on physics textbooks (104/104)
- Architecture diagram (
docs/ARCHITECTURE.drawio)
Highest-Priority Open Work (collaborators welcome)
- A semantic or LLM-backed content critic. The single highest-value component for practical coverage — the binding constraint on every headline result.
- A live production deployment where chains are written automatically at ingest and the registry gates a serving path under real query load.
- Negative controls on the physics corroboration corpus (v3) and corroboration tested on organically weak chains.
- Repeated-narrator penalties. Two passes through the same lossy transformation are strictly worse than one. Recorded but not yet penalised.
- Per-cell evidence-count monitoring as a first-class operational metric.
- Extension to probabilistic and provisional claims.
Good First Issues
- An alternative critic (sentence-transformers embedding, CrewAI integration)
- A seed-grade bootstrapper from published benchmark accuracies
- A new
CorrelationDetectorusing embedding similarity or model-card lineage - Extend semantic corroboration to multi-source corpora
- A calibrated
TransitionPolicyfrom your own pipeline's §8 experiment data - Domain-specific
ContentCriticwith formula canonicalisation - Pipeline adapters for CrewAI or AutoGen tracing
In the Wild#
Live Verify — Independent Comparison
Paul Hammant — co-creator of Selenium (~22 years ago), maintainer of trunkbaseddevelopment.com, ex-ThoughtWorks, based in Scotland. Reached out to Ali directly after finding the paper.
A dedicated comparison document: Live Verify vs. Isnād–Rijāl
Both attack the same root problem — why should I trust this claim that reached me through a chain of hands? — and both reach for a provenance answer rather than a plausibility answer. But they sit at opposite ends of the trust pipeline. They are complementary rather than competing.
| ISNAD | Live Verify | |
|---|---|---|
| Layer | Inside an AI pipeline, per-claim | Between organisations, per-artifact |
| Mechanism | Graded reputation registry + Bayesian loop | Cryptographic hash + domain endpoint |
| Trust root | Accumulated narrator track record | An authorizedBy chain to a jurisdiction's root namespace |
| Output | serve / review / quarantine | verified / not-verified + status |
| Truth stance | Probabilistic (how reliable) | Binary (unaltered vs. altered) |
| Handles | Distortion by AI transformers | Forgery/tampering of documents |
"ISNAD's README carries a 'What's Validated vs. What's Not' box and reports two of its own results as inconclusive rather than positive. … Both projects treat scoping-out over-claims as a first-class feature, not an afterthought." — from the Live Verify comparison document.
Read the full comparison → · Have your own project or comparison? Add it — reach out.
Indexing, Archiving & Aggregation
Also discoverable via NASA ADS, Google Scholar, Semantic Scholar, Connected Papers, alphaXiv, CatalyzeX, Litmaps, scite, OpenAIRE, and Hugging Face Daily Papers. Reddit launch threads — URLs pending; will be linked when recovered.
Prior & Adjacent Work
- PROV-AGENT (Souza et al., IEEE e-Science 2025) — extends W3C PROV to capture agent interactions.
- Agent-Sentry (Sequeira et al., 2026) — bounding LLM agents via execution provenance.
- HadithRank (Hawramani) — models isnād topology using information theory, treating independent chains as redundant channels. That a mathematician grading hadith and this paper grading agents converge on the same redundancy rule is evidence the mapping is structural, not decorative.
- The Provenance Paradox (Prakash, 2026) — when delegates can inflate self-reported quality, quality-based routing can perform worse than random. Consistent with ISNAD's design intuition but not a replication.
- RAPS (2026) — reputation-aware publish–subscribe with Bayesian overlays.
- TrustTrade (2026) — replaces "uniform trust" with cross-agent consistency weighting.
- AgentPoison (Chen et al., NeurIPS 2024) — red-teaming LLM agents by poisoning memory/knowledge bases; the demonstrated threat the mawḍūʿ tier defends against.
- Karpathy's "LLM Wiki" sketch (2026) — the compiled-knowledge-base system class that motivates the whole framework.
Two public repositories independently reference isnad-style provenance chains: cognalith/isnad (agent skill security with provenance verification) and VAMaksimov/threat-intel-format (threat intelligence schema with isnad-style provenance chains). Neither is affiliated.
Links#
Cite#
Paper
@article{raja2026grading,
author = {Ali Zahid Raja},
title = {Grading the Narrators: An Isnad--Rijal Framework for
Claim-Level Provenance in Multi-Agent Knowledge Systems},
year = 2026,
doi = {10.48550/arXiv.2607.24117},
eprint = {2607.24117},
archivePrefix = {arXiv},
primaryClass = {cs.AI}
}
Software
@software{raja2026isnad,
author = {Ali Zahid Raja},
title = {Isnad--Rijal Framework: Reference Implementation},
year = 2026,
doi = {10.5281/zenodo.21216873},
orcid = {0009-0003-7875-4590}
}
LLM usage disclosure: drafting and editing were assisted by a large language model; all technical claims, design decisions, experimental results, and final text were reviewed by and are the responsibility of the author.
Frequently Asked Questions#
What is ISNAD in one sentence?
ISNAD is an open-source framework that grades every agent, scraper, and model in a claim's transmission chain — adapted from classical Islamic hadith science — to tell you how much to trust a claim that has passed through multiple AI hands.
How is this different from W3C PROV or PROV-AGENT?
PROV records what happened — it captures chains, agents, and activities. ISNAD grades who did it and whether to trust the result. PROV gives you a transmission log; ISNAD tells you which link in that log is the weakest, and what action to take. See the positioning tables for a detailed comparison.
Does ISNAD outperform confidence-based gating?
The matched-coverage comparison could not be completed. With the reference content critic, ISNAD-gated serving could not be driven above 4.8% coverage, so a meaningful head-to-head comparison was not possible. ISNAD admitted zero corrupted claims among those it served — but it served very few. See the full evaluation.
What does "the chain is graded by its weakest link" mean in practice?
A ṣaḥīḥ model at the end of a chain cannot rescue a ḍaʿīf scraper at the start. The chain's grade is the minimum grade among its narrators — adjusted by transformation type (destructive vs. generative) — because information lost early in the chain cannot be recovered downstream. The chain animation demonstrates this visually.
What is mutābaʿāt / corroboration, and when does it not fire?
Mutābaʿāt is the mechanism by which independent chains asserting the same claim corroborate each other, upgrading the claim's grade (capped at ṣaḥīḥ). It does not fire when chains share a narrator, model family, or upstream source — the classical madār case. It is validated at 603/603 on Wikipedia and 104/104 on physics textbooks, with 8/8 negative controls passing. No negative controls have been run on the physics corpus yet.
Why grade narrators per domain instead of globally?
A model precise on historical dates may be unreliable on quantum-mechanical equations. The registry key is (narrator, domain), not narrator alone. This mirrors classical practice — rijāl scholars graded transmitters domain-specifically as a matter of course.
What happens when a model version changes?
A model-version bump resets the narrator to ungraded until re-evaluated. Version drift is treated as a new narrator — track record does not carry forward. This is intentional: a model at v2 may behave differently from the same model at v1, and inherited reputation would be misleading. Known blind spot: hosted models whose weights change without a version-string change will silently retain a grade they no longer merit.
What is the "unverifiable" verdict and why does it matter?
"Unverifiable" is the third column of the decision matrix — the classical position of tawaqquf (suspension of judgment). A content critic that cannot evaluate a claim must be able to say so rather than defaulting to "consistent." Routing unverifiable claims conservatively is what makes the framework safe under a weak critic. It is also the binding constraint on practical coverage — most real claims receive this verdict under the reference critic.
What did the evaluation fail to show?
Three things: (1) ISNAD cannot outperform a baseline at matched coverage because it cannot reach matched coverage — the content critic caps it at 4.8%. (2) The jarḥ–taʿdīl loop recovered 3 of 4 narrator grades but missed the highest-fault narrator because it lacked sufficient calibration evidence. (3) The confidence baseline is synthetic and proves nothing about real LLM self-confidence. These are reported in the same typographic weight as the successes. See the evaluation section.
Is this a religious claim or ruling?
No. ISNAD borrows a methodology — the epistemological architecture of isnād criticism — from classical Islamic hadith science. It makes no religious claims, issues no rulings, and does not position any computational system as an authority within the religious sciences. See the positionality statement.
Can I use this in production today?
The offline pipeline — chain grading, narrator registry, content criticism, decision routing, corroboration — has been exercised end to end on 20,000 claims. The LangChain integration is shipped and tested. A live production deployment where chains are written automatically at ingest and the registry gates a serving path under real query load is open work — it is the highest-priority next step and collaborators are welcome.
How do I cite it?
Two BibTeX entries are available in the Cite section — one for the arXiv paper and one for the software. Paper: arXiv:2607.24117 (cs.AI). Software DOI: 10.5281/zenodo.21216873.
What licence is it under?
Code: Apache-2.0. Paper and documentation: CC BY 4.0. Both are permissive open-source/open-access licences suitable for academic and commercial use.
How can I contribute?
The CONTRIBUTING.md lists good first issues — including alternative critics, seed-grade bootstrappers, new correlation detectors, and pipeline adapters. The highest-impact contribution is a semantic or LLM-backed content critic, which is the binding constraint on practical coverage. Pull requests welcome.