Then I ran the binary and found three of them were lying.

There is a repo in our org called gnosis. Steffen wrote the core of it back in February. A web scraper that turns sites into clean Markdown, with a provenance block on every file saying where the bytes came from.

Eleven commits. Then nothing for six months.

Last night I pointed Pi.dev at it. This morning it was on PyPI at v1.4.0 with a docs site, a CI matrix, WARC archival, an SSRF guard, and 128 tests.

I did not open a single file.

That is the part everyone wants to hear about. It is also the least interesting part of what happened.

What I actually typed

I want to be precise about this, because “I used AI to build it” is a sentence doing no work at all right now.

I never wrote code. I never read code. What I wrote was instructions, in five categories:

Plan. Not “build features.” I told it to settle a positioning thesis first, with an adversarial panel arguing it out. What came back was a ROADMAP with a north star, an architecture spine, a numbered issue list, and a section called “what we deliberately skip.” That last section is the one that mattered. A plan that only says yes is a wish list.

Audit. After every batch, a QA pass against the work it had just done. You can see it in the log. feat: enforce robots.txt at 01:55, then fix: robots crawl-delay parsing + fail-open at 02:09. It found its own bug in fourteen minutes.

Research. Where to steal from, by name. The ROADMAP has a table with a source column: WARC and replay from warcio and pywb, robots enforcement from scrapy, the plugin-extras pattern from markitdown, chunking from crawl4ai. Every capability tagged with where the idea came from.

Scope. No agentic browsing. No V8 embed. No LLM extraction in core. Saying no is the highest-leverage instruction you can give a harness, because it will happily build everything you don’t forbid.

Ship. Version bumps, changelog, GitHub Pages, PyPI. It pushed and published. I approved.

The log is the receipt:

Feb to July, human-written — 11 commits

Sept 1, 18:58 to Sept 2, 14:50 — 37 commits

PyPI — 1.1.2 to 1.4.0

Package — 4,123 lines, 128 tests, 74% coverage

Twenty hours. Zero files opened.

The marketing was the same loop

This is the half nobody writes about. Everyone treats agents as a code tool. The code was the easy part.

I gave it one instruction: figure out where this can win, then write everything as if that were true.

It came back with a positioning thesis I would have taken a week to reach on my own:

Everyone in this space competes on cleanliness, scale, and agents. Nobody owns verifiability.

So gnosis became “the auditable one.” Markdown is a projection. The raw bytes and the WARC are the source of truth. Once that sentence existed, every other decision fell out of it: hash the bytes, not the derived text; archive to a format that replays; put the verification command in the README so a reader can check the claim themselves in one line.

From that thesis it wrote a README with the comparison table, the keyword set in the package metadata, an llms.txt so other people's agents can read the project, a MkDocs site, a demo GIF, CONTRIBUTING, a security policy. All of it consistent, because all of it came from one settled position instead of five separate sessions of me guessing.

That is the real unlock, and it has nothing to do with typing speed. A harness will hold a thesis across forty artifacts without drifting. I won’t. By file thirty I’d be writing a slightly different pitch than I wrote in file one.

Then I audited it properly

Twenty hours of clean commits. Tests green. Lint clean. Docs live.

So I ran the binary against a page I controlled and checked the output against what I knew was in it.

Resume didn’t resume. Crawl capped at two pages, then re-run with the cap raised. It found nothing. The code skipped already-seen URLs before extracting their links, so the frontier collapsed and the crawl exited immediately. There is a commit titled fix: resume skips refetched URLs. It did not fix resume.

The archive ate itself. The WARC file opened in write mode instead of append. Two fetches into the same folder, two Markdown files, two blobs in the content store, and a WARC containing only the second one. The headline feature of the whole positioning, quietly deleting its own evidence on every re-run.

The fidelity metric reported the opposite of the truth. I fed it a page with a data table and two content divs whose CSS classes contained ordinary words like share and cookie. The boilerplate stripper deleted all three. Silently. The file it produced carried this:

retention_ratio: 1.0613

I measured it by hand. 372 characters of visible text in, 123 deleted. True retention was 0.67. The number was inflated because it compared Markdown character count against source text length, and Markdown syntax plus expanded absolute URLs padded the numerator enough to hide a third of the document going missing.

An audit field that certifies a lossy capture as complete is worse than no audit field. That one is on me for asking for a metric without asking what it should read when things go wrong.

And then the one that actually stings. I asked for per-host rate limiting. It keyed the timer on the full URL instead of the hostname. A crawler never fetches the same URL twice, so the timer never fired.

Rate limiting was gone. Robots.txt Crawl-delay went with it, because it feeds the same dead gate. A tool whose entire pitch is polite, auditable crawling, hammering targets as fast as the event loop allowed, while the README promised otherwise.

The pattern is the point

Look at what it got right and what it got wrong, and it stops looking random.

It nailed the DNS rebinding fix. Genuinely good work: pin resolution, validate every address in the answer set, fail closed if the set mixes public and private, keep SNI intact so certificate verification isn’t weakened. Better than most production code I’ve read.

It nailed everything with a clean specification. Append instead of truncate. Add a token count. Publish to PyPI.

It failed every single thing that required looking at the output and asking whether it was absurd.

Nothing in the loop ever ran the binary and read the bytes. The tests were green because the tests asserted that the code does what the code does. Not one of them compared extracted text against source text. Not one asserted that five pages at 2000ms should take longer than eight seconds. Both are five-line tests. Neither existed, because I never asked for them, and a harness does not volunteer doubt.

The harness closes the implementation gap and widens the verification gap.

That is the whole finding. Generation stopped being my bottleneck twenty hours ago. Verification became the entire job, and I had not moved a single resource there.

What I changed

Four rules, all of them cheap:

  • The adversarial check ships before the feature. For every claim in the README, name the test that could catch it lying. Write that first.
  • Every metric must be able to report failure. If I can’t describe the input that makes the number look bad, the number is decoration.
  • Run the binary against something you already know the answer to. Not the test suite. The actual artifact, against a page you wrote yourself.
  • Treat “fixed” as a hypothesis. Two of my three findings survived a round of commits aimed directly at them. A commit message is a claim, not a result.

I still think zero-touch is real. Twenty hours, a dead repo to a published package, and the security work is better than what I’d have written by hand at 2am.

But the version of this story going around right now stops at the commit count. That version is a lie of omission, and the people repeating it are going to ship the rate-limiting bug into somebody’s production crawler and never know.

The machine will build anything you name. It will not tell you that 1.0613 is an absurd number to print on a page it just gutted.

That part is still yours!

Code: pip install gnosis-markdown · github.com/SHCV-it/gnosis · MIT