How a Solo Developer from Pakistan Landed on LangChain’s Middleware Integrations List — Alongside OpenAI and Anthropic
A pull request, a 1,200-year-old methodology, and what the milestone actually means

On 29 August I opened a pull request against langchain-ai/docs. One row in one YAML file. The repo had 94 open pull requests and 280 open issues at the time, and LangChain had introduced a new submission process for external integrations barely a month earlier. I assumed it would sit there for weeks.
But it merged on 31 August, just two days later!
What surprised me wasn’t the speed. It was that the maintainer who merged it, Naomi Pentrel, didn’t bounce it back with a checklist. She pushed three more commits herself, force-pushed my branch, and finished the listing: registered the package, added a provider card, added a row to the middleware downloads table. Then she approved and merged
ISNAD is now in LangChain’s official documentation, in the All middleware table, sitting in a list that includes OpenAI, Anthropic, AWS and Microsoft Foundry.
I want to write about how that happened. But I want to be precise about what it means, because the interesting part isn’t the listing.
The thing nobody’s measuring
Here’s the problem I couldn’t stop thinking about.
When your multi-agent system hands you an answer, that claim has passed through several hands. A source document. A scraper. An ingestion model that pulled structure out of it. A synthesis model that wrote the final sentence. Every one of those hands can drop something, distort something, or invent something outright.
Now look at what we measure. Confidence scores tell you how certain the last model was. That’s it. The last model has no idea that the scraper three hops back mangled a table, or that the extraction step hallucinated a date. Its confidence is high anyway. That’s not a bug in the confidence score, it’s the confidence score measuring the wrong object.
Provenance tooling helps, but only partly. OpenTelemetry-style tracing records what happened. That’s a log. A log is necessary and it is not a judgement. Nothing in it tells you this answer is untrustworthy because of something four steps upstream.
A confident model at the end of a chain cannot repair a corrupted extraction at the start of it. Once I saw the problem in those terms, I couldn’t unsee it — and I also recognised it, because I’d seen the solution before, in a completely different context.
Where the answer came from
I’ve spent years building Islam & AI , a platform now serving over 143,000 users across 150+ countries. Working with that corpus meant working with hadith — and hadith scholarship spent twelve centuries on a structurally identical problem.
The question they faced: do you trust a statement transmitted through a chain of human narrators, none of whom you can interview? Their answer had four moving parts that map almost directly onto multi-agent AI:
Isnād — every claim carries its complete chain of transmission. Not optional metadata. A claim without a chain is a different class of object entirely.
Rijāl — every narrator has a documented, contestable grade. Not reputation by vibes. Recorded criteria, applied consistently, revisable when new evidence arrives.
The weakest link governs. A chain is graded by its worst transmitter, not its average and not its best. No downstream reputation repairs an upstream fabricator.
Matn criticism runs separately. The content gets criticised independently of the chain. A sound chain carrying contradicted content isn’t an error, it’s the most informative signal in the system.
Substitute “narrator” with “agent, model version, or scraper” and that is a design specification for claim-level trust in multi-agent AI. That’s the whole thesis of the paper I published on arXiv in July (2607.24117), and of the library that came out of it.
Building it — and trying to break it
ISNAD is the implementation. pip install isnad. npm i insand. Every claim carries its chain, every transmitter keeps a living per-domain grade, the chain caps at its weakest link, and the result routes through a decision matrix to a concrete action: serve, serve with caveat, hold for corroboration, or quarantine and flag the narrator. The grading path is deterministic, no LLM sits inside it, so it runs entirely local with no API calls.
The obvious objection to all of this is that weakest-link aggregation is arbitrary. Why the minimum? Why not a weighted average? It feels right, which is exactly the kind of reasoning I don’t trust.
So I tried to falsify it. Classical hadith scholarship left behind something unusual: a large corpus of transmission chains where human experts already recorded their verdicts, produced entirely independently of anything I built. I graded 577,024 real chains with the library’s rule and compared against the scholars’ own judgements.
Cohen’s κ of 0.871 strict, 0.761 lenient.
For context, inter-scholar agreement on the same chains is κ 0.331. The rule agrees with the tradition more than the tradition agrees with itself.
I want to be careful here because that number is easy to oversell. It is not evidence the rule is smart. It’s evidence the rule is consistent where humans were noisy, which is genuinely useful if you’re building a machine system, and is a weaker claim than the headline figure suggests. The strict/lenient gap is also where the interesting failure modes live. I report the inconclusive parts of the evaluation at the same weight as the wins, because a framework about trust that hides its own weak results has failed its first test.
What the listing actually is
Now back to the pull request, and the part where I have to be honest with you.
LangChain’s page states plainly that community integrations are contributed on an open-source basis and are not managed or maintained by LangChain. So let me be exact about the claim:
- ✅ ISNAD is listed in LangChain’s official middleware integration documentation.
- ❌ It is not an “official LangChain integration.” It is not a partnership. Nobody at LangChain has endorsed it.
The table is sorted by download volume, and the sort value updates on every docs build. Today ISNAD sits sixth. Next month it might not. Anyone can open the same pull request I did.

So what is it worth? A docs listing is a credibility marker, not a demand signal. Nobody is going to adopt a provenance framework because it appeared in a table. What it does is remove one argument from every conversation I was already having, I no longer have to establish that this is a real, working, used project before I get to talk about the idea. That’s not nothing. It’s also not the thing.
The thing is the κ 0.331 comparison. The listing is downstream of it.
One detail I did enjoy: the description I submitted made it through verbatim, including the word mawḍūʿ — the classical term for a fabricated chain. There was no English equivalent worth substituting, so I didn’t substitute one. That term now ships in LangChain’s documentation.
Talent is global. Opportunity is not.
I write that line on everything, so I should say what I mean by it here.
I did this from Wah, Pakistan. No lab, no affiliation, no institutional email, and no co-author. The pull request was reviewed on its contents, by someone who had never heard of me, in a repository where nobody could see where I was sitting. That is the part of open source that actually works, and it is worth naming, because most systems aren’t like that.
The methodology isn’t mine either. The rigour belongs to twelve centuries of muḥaddithūn who built a verification system so careful it still outperforms its own practitioners’ consistency. The transfer is my contribution. That’s the whole of it, and it’s enough.
If you want to look
- Code: github.com/alizahidraja/isnad — pip install isnad .npm i isnad
- Paper: arXiv 2607.24117
- Listing: https://docs.langchain.com/oss/python/integrations/middleware
Issues and disagreement welcome. Several of the open issues in that repo came from people who turned up specifically to argue with me, and the project is better for every one of them. If you think the weakest-link rule is wrong, I’d rather hear it now than find out in production.