The benchmark could have failed. That was the point of building it.

ISNAD had 545 tests and 96% coverage. A paper on arXiv. A PyPI package. Strangers filing good bug reports.

What it did not have was a single number showing the thing actually works.

The whole pitch is that you can grade machine claims the way hadith scholars graded testimony — chain of transmitters, reliability grade per transmitter, weakest link decides. Nice pitch. Twelve centuries of pedigree but completely unmeasured.

A trust framework with no measured validation is a claim, not a fact. Mine had a credibility surplus in code and a deficit in evidence.

So I went and measured it.

The question had to be one I could lose

The failure mode with benchmarking your own project is obvious. You design an evaluation you’ll pass, run it, and present the result like it fell out of the sky.

So I picked a question with no soft landing:

Give ISNAD the same narrator grades the real scholars assigned. Does its weakest-link rule reach the same verdicts the scholars reached?

One scalar out the other end. No “well, in a certain light.” Either the rule reproduces the tradition or it doesn’t.

The good part: the answers already exist. Written down. By people who had no idea a Python library would one day be checked against them.

Where the ground truth comes from

The benchmark runs against emadjumaah/hadith-kg, CC-BY-4.0:

  • 715,790 hadiths
  • 577,024 chains
  • 49,844 narrators
  • 127,863 criticism statements
  • 1,015 named critics

Narrator grades come from Ibn Ḥajar’s Taqrīb al-Tahdhīb — the standard twelve-tier ranking, thiqa down to kadhdhāb.

Chain verdicts come from the scholars’ own words. Phrases like إسناده متصل رجاله ثقات — chain connected, narrators reliable.

Nothing synthetic or nothing I generated. The dataset is pinned by SHA-256 in the repo so anyone running it gets exactly what I ran.

Four rules so I couldn’t cheat

Preregister the mapping. The Arabic-grade → ISNAD-grade translation went into the repo before a single result was computed. If I’d tuned it afterward to chase a nicer number, the commit order would show it.

Use Cohen’s κ, not accuracy. Ṣaḥīḥ dominates the corpus. Answer “ṣaḥīḥ” to everything and you post a great accuracy score having learned nothing. κ discounts the agreement you’d get by chance, which makes the lazy strategy worthless.

Run negative controls. Shuffle the narrator grades: κ collapses to 0.047. Majority-class baseline: 0.000. If the signal were an artifact of how I’d wired the pipeline rather than of the grades themselves, the shuffled control would still score. It doesn’t.

Bucket every disagreement. The residual isn’t buried in one number. Each class of mismatch is named and counted — mapping ambiguity, continuity gaps, severity convention, corroboration dependence. Some of those are my problem. Some are the tradition disagreeing with itself. Saying which is which is most of the actual work.

The numbers

Cohen’s κ ISNAD vs scholarly consensus (strict default) 0.871 ISNAD vs consensus (lenient opt-in) 0.761 One scholar vs consensus 0.450 Scholars vs scholars 0.331

Read the bottom row again.

The number that matters is 0.331, not 0.871

I could write the headline “ISNAD beats 1,200 years of scholarship.” It would travel but it would also be nonsense.

The scholars agree with each other at κ = 0.33. The ground truth is contested. Hadith criticism is an argument, not a lookup table someone forgot to digitise. A library that outscored the entire tradition would be a library with a bug in its evaluation harness.

What 0.871 says is smaller and more useful: ISNAD is a faithful, deterministic implementation of the consensus. Same inputs, same verdict the tradition converges on, reproducibly, at scale, no human in the loop.

That’s what a library should be. Nothing more heroic than that.

Three things the benchmark taught me

I built it to validate. It ended up teaching things which I didn’t expect.

1. My biggest error bucket has a classical name

The largest remaining disagreement was chains where the scholars wrote some version of weak alone, but ḥasan if corroborated. ISNAD said weak. They said acceptable. On paper, a miss.

So I checked whether those chains actually have independent corroborating routes elsewhere in the corpus.

88–92% of them do.

They weren’t being inconsistent. They were applying mutābaʿa — independent parallel transmission strengthens a weak report. ISNAD already implements that. The benchmark validated my corroboration engine from the outside, using data that knows nothing about my code.

You can’t manufacture that result ¯\_(ツ)_/¯

2. Reliability has a timeline, and they already knew

161 narrators in the corpus declined over their lifetime — age, illness. Their transmissions touch 35% of all chains.

The scholars didn’t throw those narrators out. They dated the decline. Reliable before, unreliable after. Same person, two grades, split by a year.

ISNAD has period-sliced grading — get_grade_as_of(). I built it because it seemed principled. Turns out it's load-bearing across a third of the corpus.

The modern version writes itself. A model, an API, a data source is not permanently trustworthy or permanently suspect. It has a reliability timeline. Evidence gathered before a drift event shouldn’t be invalidated by the drift, and evidence gathered after shouldn’t coast on the old reputation.

3. My default was wrong

This one cost me something.

When a narrator had no grade at all — nobody wrote a criticism statement about them — ISNAD treated them as passable. Unknown, give them the benefit of the doubt. Felt generous. Felt reasonable.

κ = 0.761.

The classical position on an unknown transmitter, majhūl, is the opposite. Unknown means you can’t vouch for them, so it caps at ḍaʿīf. I flipped the default to match.

κ = 0.871.

Eleven points of kappa, and the 1,200-year-old method was right and my intuition was wrong. The old behaviour is still there as lenient_unknown=True — opt-in, documented, not quietly deleted and not quietly the default.

If you want the short version of why I keep mining this tradition: that. They converged on conservative defaults under adversarial pressure. I converged on a generous default because it felt nice. One of us had been tested.

What shipped in 2.6.0

  • ISNAD-Bench (bench/) — self-contained and reproducible, four measurements: chain agreement, corroboration ablation, human ceiling, ikhtilāṭ.
  • Strict default — ungraded narrators cap at ḍaʿīf. Lenient behaviour is opt-in.
  • py.typed, --json export, README section.
  • Multi-provider LLM critic, carried over from 2.5.0.

597 tests pass. ruff and mypy --strict clean. Every change went through the standing multi-persona audit before it was allowed in.

What it doesn’t cover: the benchmark measures the deterministic grading rule — weakest link, corroboration, period slicing. The LLM-facing side is a separate evaluation problem I haven’t solved yet.

The bit that isn’t about hadith

Most systems answer “how much should I trust this?” with a confidence score. One opaque float. Unauditable, uncomparable across systems, indefensible the moment someone asks why 0.83.

The alternative isn’t a better float. It’s carrying the chain, grading the sources by named criteria, applying a conservative combination rule, and quarantining instead of guessing.

That was always the argument. As of this release it has a number attached, produced by a preregistered protocol with negative controls, against ground truth nobody can accuse me of writing.

The diagnosis was the product is ahead of its evidence.

It isn’t anymore.

ISNAD 2.6.0 is on PyPI. The benchmark, the preregistered mapping, the negative controls and the full error buckets are all in the repo — including the disagreements. If you can break the result, please do. That’s what the pinned dataset hash is for.