The hadith scholars solved AI’s hallucination problem in the ninth century. I went back and read their notes.

You ask a language model for a narration about, say, the merit of seeking knowledge. It gives you one. It’s in Arabic. It’s stylistically perfect — the diction is right, the rhythm is right, it opens the way these narrations open. It attributes it to Bukhari with a number. And it does not exist.

Not “is weakly attested.” Does not exist. The model produced the shape of a hadith, because it has seen ten thousand of them and shape is what these systems learn.

I want to be precise about why this is different from ordinary hallucination. If a model invents a Python function, your code fails and you find out in eleven seconds. If a model invents a citation in an essay, a reviewer checks and you’re embarrassed. But a fabricated narration in Arabic, correctly formatted, attributed to a real collection, is portable. Somebody screenshots it. It goes into a WhatsApp group. It gets repeated in a khutbah by someone who trusted the screenshot. There is no error message. The failure propagates silently through exactly the channel the material was designed for.

And here is the part that should bother anyone building in this space: the more capable the model gets, the better the fabrications become. Fluency is not a proxy for truth. It’s a proxy for how convincing the falsehood will be.

The tradition already solved this

The thing that reframed the whole project for me is that this is not a new problem. It’s a very old problem, and it was solved comprehensively by people who had no computers and enormous discipline.

In the first two centuries of Islam, a large volume of fabricated narrations entered circulation. The reasons varied — political factions manufacturing support, sincere storytellers improving a good line, sectarian partisans backfilling their positions. The material was in Arabic, it was well-formed, it was attributed to real authorities. Structurally: exactly the problem I just described.

The response was not to build a better content filter. Nobody tried to determine truth by examining the text alone. Instead the scholars did something that, in retrospect, looks remarkably like engineering.

They made transmission the unit of verification rather than content.

Every narration had to carry its chain — isnād — the list of people who passed it from the Prophet ﷺ forward, each one named. And then they built a second discipline, ʿilm al-rijāl, the science of the men: a biographical apparatus covering tens of thousands of transmitters. When did each person live, who did they demonstrably meet, what was their memory like, did they ever contradict themselves, did anyone credible ever call them careless. Multi-volume works, cross-referenced, maintained across generations.

A narration was then graded on the weakest link in its chain. Not the average. Not the majority. The weakest. One transmitter with a documented memory problem caps the grade of the entire chain regardless of how impeccable everyone else was, and regardless of how beautiful or how theologically convenient the text is.

Sit with that design decision for a second. It’s a conservative propagation rule, applied to a directed chain, with per-node reliability metadata, where uncertainty dominates rather than averages out. If you handed that spec to a distributed systems engineer today they would recognise it immediately.

Porting the method

I spent a good part of last year formalising this. The result is ISNAD — a claim-level provenance framework for multi-agent AI systems, where each claim carries the chain of agents, tools, and retrievals that produced it, each link gets a reliability grade, and the claim inherits the weakest link’s grade. Unverifiable steps are flagged rather than smoothed over.

I’m not going to pretend I invented anything. I read the method and translated it.

What I did test is whether the translation survives contact with reality. The library now ships a benchmark that runs the weakest-link rule against 577,024 real hadith chains and compares its output to the gradings the classical scholars actually assigned. Agreement came out at Cohen’s κ of 0.871 under strict settings. For calibration: inter-scholar agreement — how much the human scholars agree with each other — sits around κ 0.331.

That gap is worth being honest about, in both directions. It does not mean the algorithm is better than the scholars. Scholars disagree because they’re weighing genuine judgement calls, contested biographical reports, and considerations the rule doesn’t model. What it means is narrower and still useful: the mechanical part of the method is mechanical enough to be reproduced, and the scholars applied it far more consistently than they get credit for.

What this looks like in a product

islamandai.com is the applied version, and the design is mostly a list of refusals.

It does not generate Arabic scripture. It will not produce the Arabic of a verse or a narration as model output. It gives you the reference — surah and ayah, collection and number — and links to Quran.com and Sunnah.com. The canonical text comes from a canonical source, not from a probability distribution. This costs some convenience. It is not negotiable.

It does not issue rulings. This is the one people push back on, so let me take it seriously.

There is a published ruling from IslamQA that it is not permissible for an ordinary Muslim to seek a fatwa from an AI system or act on its answer. The reasoning is sound: the system isn’t qualified, it doesn’t understand the questioner’s circumstances, its sources may be unknown, and fatwa has always been sensitive to the particular person asking. The same ruling permits something else, though — a researcher may use AI to gather material and understand a landscape, on condition that he verifies the sources.

That condition is the entire product. Islam & AI is built to sit inside that carve-out and nowhere else. Ask it whether your specific divorce was valid and it will decline and tell you to see a scholar. Ask it what the four schools hold on a question and where their disagreement originates and it will lay that out with references you can check.

I find it clarifying rather than limiting that the scholars drew this line before I got here. My job was to build something that respects it, not to argue with it.

It does not resolve disagreement by hiding it. A recent survey paper on Islamic chatbots proposes three axes for comparing these systems: authority posture, technical grounding, and sectarian encoding. The third is the one most tools handle badly — they pick a position and present it as the answer, and the user never learns that a question was contested at all. Flattening eleven centuries of scholarly disagreement into a confident paragraph is its own kind of fabrication. When there’s a real difference, the tool shows the difference.

It says when it doesn’t know. Which turns out to be the hardest engineering problem in the entire build, and the one I’m least finished with.

What I’m not claiming

I’d rather state the limits than have someone find them.

This is not a scholar and cannot become one. It has no knowledge of your circumstances, which is precisely what fatwa depends on. Its retrieval layer can surface the wrong passage. Its summaries of scholarly positions compress things that shouldn’t always be compressed. The refusal behaviour is enforced through architecture and prompting, and while I’ve tested it hard, I will not stand here and tell you it’s impossible to defeat — anyone who tells you their LLM product is impossible to defeat is either lying or hasn’t tried.

What I will claim is narrower: it’s built so that when it fails, you can catch it. Every substantive claim points somewhere you can check. That’s the whole design.

The tradition’s insight wasn’t that transmitters were trustworthy. It was the opposite — it assumed they weren’t, and built an apparatus that made unreliability visible.

That’s a good specification for an AI system. It might be the only one that works.

islamandai.com is free, runs in any language, and is funded by donations. It isn’t for sale.

If you use it, please try to break it first. Ask for something obscure, something contested, something that doesn’t exist. If you get a fabrication out of it, post the screenshot publicly and tag me. That’s more useful to this project than any amount of goodwill.