What actually happened when I stopped hiring people and started running agents.

I’ve been saying “AI agents changed how I work” for a year now. Nobody really believed me, and to be fair the proof was thin. A couple of demos. Some repos. Nothing with a card on file.

So this time I kept receipts.

Below is everything that went from nothing, or from dead, to live in production with real users or real money attached, between the last week of July and today. I did it alone, from Wah Cantt, Pakistan. My total tooling spend for the whole run was less than what a junior dev costs for a week.

Two things first so nobody accuses me of inflating. Tianlu and Gnosis are SHCV products, and I built them with Steffen Höhne, who wrote the original core of both. Everything else is mine and “live” means live: a URL you can open, a pip install you can run, a restaurant you can order from tonight.

Here’s the list. Then I’ll tell you exactly how.

1. ISNAD went from a blank file to a paper, a library, and a company

In July I had an idea and nothing else. Classical hadith science has a 1,200-year-old method for grading who transmitted a claim, not just what the claim says. Multi-agent AI pipelines have the exact same problem and no method at all. A scraper hands a fact to a model, the model hands it to another model, and nobody grades the hands.

What exists now:

The paper. “Grading the Narrators” is on arXiv (2607.24117), on Zenodo, on Hugging Face Papers. My first paper!

The library. pip install isnad. Version 1.0.0 shipped on 7 July. As I write this it's at 2.23.0. That's roughly sixty-five releases in two months, all from one keyboard. It has LangChain, CrewAI, LlamaIndex and OpenTelemetry integrations, an MCP server, a JavaScript verifier on npm, and it's listed in LangChain's official docs as a community middleware. The PR was merged by a LangChain maintainer

The benchmark. This is the one I’m proudest of. I ran ISNAD’s weakest-link rule against 575,060 real hadith chains that scholars graded over the last twelve centuries. Cohen’s κ against scholarly consensus: 0.871. For context, the scholars agree with each other at 0.331 on the underlying narrator grades. I didn’t build a benchmark I could pass. I built one I could fail in public, ran it, and then flipped the library’s default to strict because the data said so.

The company. isnad.islamandai.com is live. Hosted ISNAD, an API key, a dashboard, a card on file. The core library stays Apache-2.0 forever, that’s written into the repo. The hosted surface is what pays for the next twelve months of me working on it.

Paul Hammant, who co-created Selenium, wrote an unprompted comparison between his Live Verify system and ISNAD. So I built the integration. That’s how most of the good things in this list happened. Someone credible poked at it, and I shipped the answer instead of a reply.

2. Islam & AI was dead for a year. It’s back, and it pays for itself (hopefully!)

I launched islamandai.com in 2023. It reached about 145,000 people across 190+ countries, free, running on donations, and then I let it rot. The UK company behind it got dissolved. The backend was on a model that no longer made sense and wasn’t even supported

In August I rebuilt the whole thing. New backend on DeepSeek. New frontend. Usage-cost tracking per query so I actually know what it costs to run (about $30 to $45 a month at current traffic). It references Quran verses and hadith by number and points you to Quran.com and the collections instead of generating Arabic text it can’t stand behind.

Then I added a way to pay for it. Free core chat stays free forever. Lite is $20 a year, Pro is $50, Benefactor is $100. Paddle rejected the domain because “AI chatbot” is a banned category for them. I appealed, moved to a different merchant of record, and kept going.

There’s an API and a dashboard at api.islamandai.com now. And the iOS and Android apps, built in Expo, are sitting in store review as I type this.

Three years of guilt, cleared in three weeks.

3. Tianlu: a knowledge base that maintains itself

This one was for a client, through SHCV. Steffen had prototyped an idea built on Karpathy’s “LLM Wiki” pattern: instead of doing retrieval per query like every RAG system, an LLM continuously compiles sources into a linked wiki, resolves contradictions at ingest time, and serves cited answers from the compiled artifact. Markdown and Git are the source of truth. Postgres is just a rebuildable projection.

The prototype had the scaffolding but the brain was never plugged in. We rebuilt v1 from scratch. The numbers from that sprint: 65 commits, 455 files changed, 69,503 lines added, 96 tests, 56 wiki pages compiled from real ingest runs. Fed it three undergraduate physics textbooks and it found 19 real cross-framework contradictions on its own. Not hallucinated ones. Actual places where two authors disagree.

Then we shipped it into Microsoft Teams which was a whole other feat (I hate azure)

ISNAD was born inside Tianlu. I needed to know how much to trust each agent in that pipeline, and there was no tool for it. So the client project produced the research project, which produced the company. That’s the loop I want you to notice.

4. Gnosis: the scraper the EU AI Office wrote back about

Gnosis (github.com/SHCV-it/gnosis) turns websites into clean Markdown with a provenance block on every file. Byte-level SHA-256 of the raw response, WARC archival, Ed25519 signed records, ai.txt consent recording, an SSRF guard. Firecrawl and Crawl4AI win on speed and hosting. Gnosis wins on the one thing none of them ship: you can prove where the document came from.

I took it from 1.1 to 2.1 in about five weeks, almost entirely by directing agents. And I want to be honest about what that looked like, because it’s the whole point of this article.

The agent’s “completeness” metric reported 106% retention on a document that had lost a third of its text. Every test was green. The rate limiter was keyed on full URL instead of hostname, so it was doing nothing. Unsigned verification returned exit 0 on tampered files. I found all of that by running adversarial audits against the agent’s own work, fixing, re-auditing, and repeating. MULTIPLE rounds. Each round found new bugs the previous fix had introduced.

That audit became a public Capture Record Specification. I sent it to the EU AI Office. They replied and pointed me at CEN/CENELEC’s prEN 18284 work on dataset governance. SHCV is now an AI Pact signatory, we submitted evidence to the UK’s DSIT consultation, and I’m in the process of joining the Estonian mirror committee for the AI standards work. From a scraper.

5. BunBites.pk: a real restaurant runs on it

My friends run BunBites in Wah Cantt. Shawarma, roll paratha, BBQ, a midnight pizza menu until 4am

bunbites.pk is live. Online menu, WhatsApp ordering, and behind it a full POS with a kitchen display, inventory, and a dashboard. Built on URY over Frappe, self-hosted, Cloudflare in front. First month costs the restaurant zero. I did the data load and setup myself (my agents did it)

This is the design partner for a white-labelled restaurant platform I’m going to sell to every restaurant in Wah, then outward. Web only. No native app. Each restaurant gets its own repo and its own domain. The template for a new restaurant gets generated by an agent.

6. ISO 42001 is in the pipeline

I’m sitting the PECB ISO/IEC 42001 Lead Auditor exam in November. Not for the badge. Because every product above produces audit evidence, and the people who charge the most money in this industry right now are the ones who can say what evidence is sufficient. I pitched an EU AI Act and ISO 42001 practice to SHCV this week. NeuroMark, Gnosis, Tianlu and ISNAD map onto specific articles. That’s the next revenue line.

7. And I’m teaching all of it, free

I ran a live session reviving my own QURAN-NLP repo (143 stars, hadn’t been runnable in years) using nothing but an agent harness, and put the recording on YouTube: “My AI Coding Agent Setup: pi.dev, Warp, and a Real Repo.” The agent gets things wrong on camera. That’s the useful part.

Now for the `How?` The actual playbook.

People read lists like this and assume either a team or a lie. It’s neither. It’s a method, and it’s cheap enough that you can start it today.

The harness is pi.dev. Minimal agent harness, four core tools, MIT, extension-first. I run it inside Warp. I have 24 extensions, 11 skills and 6 prompts on top of it. The important thing isn’t pi specifically. It’s that a minimal harness makes the context flow visible, and context discipline is the entire game on real codebases. The method ports to Claude Code, Codex, whatever comes next.

Models through OpenRouter. There are free models on OpenRouter that are good enough for most of the work. When I pay, I pay for the audit passes and the hard architecture decisions. Ten dollars gets you a full chatbot app with memory and persistence and proper security. I’m not exaggerating for effect. Islam & AI’s entire monthly compute is a takeaway meal.

Spec first, then run agents in parallel. I write the spec like I’m handing it to a contractor I’ll never talk to again. Then I run multiple agents across multiple repos at the same time. One is doing infra, one is doing the WhatsApp flows, one is writing tests. While they work I do the thinking they can’t.

Audit the agent like it’s lying to you. Because it is, a little. Not on purpose. It writes code that compiles, matches the style guide, and carries a logic regression. So after every sprint I run an adversarial pass: clone fresh, install, reproduce exact inputs, compare to known outputs. The 106% retention bug came from that. So did the rate limiter. So did the tampered-file bug. You do not skip this step. This is the step.

Publish the honesty box. Every one of these projects has a section in the README that says what’s validated, what’s partial, and what’s useless. ISNAD’s says outright that model self-confidence is uncorrelated with defects and confidence-gating doesn’t work. People trust the project more because of that section, not less. Credible people engage with things that admit their limits. That’s how Paul Hammant showed up. That’s how the AI Office replied.

Security is not optional, and agents will skip it. SSRF guards. Secrets only from environment variables. Signed audit records. PII redaction hooks. The agent will not add any of this unless you make it part of the spec and part of the audit. If you’re building a chatbot for ten dollars, the security is what separates a toy from a product. Don’t cut it.

Release constantly. Sixty-five releases of ISNAD in two months sounds unhinged. It’s the opposite. Small releases mean small blast radius, fast feedback, and a public changelog that shows the work. Nine GitHub issues came from one Reddit thread and every one of them made the library better.

Let one thing feed the next. Client project → research problem → paper → library → benchmark → company → compliance practice → teaching. None of it was planned as a ladder. I just refused to let any piece of work end as a one-off.

What it cost, and what it didn’t

Money: under a hundred dollars a month across everything, including hosting.

Time: all of it. I’m not going to pretend this was 4-hour workdays. It was long days, and I’ll write separately about what that does to you, because the honest version isn’t inspirational.

What it didn’t cost: a co-founder for every idea, six months of runway for each, and permission from anybody.

Why I’m posting this

Not to sell a course. Everything I teach is free and stays free; the only thing I charge for is client engagements.

I’m posting it because a few months ago I was a guy in Wah with a laptop and a lot of half-finished repos, and the thing that changed wasn’t the models. The models were already good enough. What changed was that I stopped treating the agent as a coder and started treating it as a team I had to manage, spec, and audit. Once I did that, the constraint became my own clarity, not my headcount.

If you have a dead repo, a half-idea, or a product you’ve been “planning” for a year, you have everything you need right now. Ten dollars and a spec.

Go ship something this week. Tag me when it’s live. I’ll read it.

Ali Zahid Raja builds provenance and trust infrastructure for AI systems from Wah, Pakistan. ISNAD: alizahidraja.com/isnad · isnad.islamandai.com. Islam & AI: islamandai.com. Gnosis: github.com/SHCV-it/gnosis. Everything else: alizahidraja.com