Everyone moved the file to next year. Here’s what actually happened, what’s still live right now, and why the standards gap is the real story.

Most AI compliance content is focused on the wrong thing.
Not “which framework applies to us.” Not “which consultancy do we hire.” Not even “when is the deadline.”
The real question is: can you prove what your system did?
That’s it. That’s the whole thing. Everything else is paperwork around that one question.
And in July, a lot of people decided they didn’t have to answer it yet.
What actually happened
The Digital Omnibus on AI — Regulation (EU) 2026/1744 — entered into force on 27 July 2026.
Parliament approved it 16 June. Council adopted it 29 June. Signed 8 July. In force 27 July.
What moved:
- High-risk obligations for Annex III systems → 2 December 2027
- High-risk AI embedded in regulated products (Annex I) → 2 August 2028
- Deadline for member states to set up regulatory sandboxes → 2 August 2027
What did not move:
- Article 50 transparency obligations → applicable since 2 August 2026
- GPAI provider obligations → in force since August 2025
- Prohibited practices and AI literacy → in force since February 2025
Read that second list again.
Article 50 is not a high-risk obligation. It’s a separate transparency layer that applies regardless of risk class. The Omnibus deliberately left it alone.
If your product talks to users, generates text, images, audio or video, or infers emotions or biometrics — those duties apply to you today. Not in 2027. Today.
Boards heard “delayed” and assumed the whole machine got pushed back. Most of the high-risk machine did. Article 50 didn’t.
The deadline nobody is talking about
This is the one I’d worry about, and I’ve barely seen anyone write about it.
Article 50(2) — machine-readable marking for generative systems — got a narrow transition. If your generative system was already on the market before 2 August 2026, you have until 2 December 2026 to have marking and detection working properly.
Four months. Not the six the Commission originally proposed.
That’s about fourteen weeks from now.
And fourteen weeks is not a lot of time to build marking and provenance metadata that survives being copied, compressed, re-encoded and re-uploaded across platforms.
That’s not a policy you write. That’s engineering.
Why it got delayed — and why that’s the actual story
Here’s the part worth sitting with.
The AI Act was never designed to be complied with directly. Article 40 set it up so that harmonised standards, once published in the Official Journal, give you a presumption of conformity. That’s the intended route. Meet the standard, you’re presumed compliant.
The standards aren’t finished.
The Commission also missed its own Article 6 guidance deadline back in February 2026.
So the technical specifications for logging, record-keeping and traceability — the specs that tell you how to prove what your AI system actually did — are still being drafted.
Everyone read that as breathing room. I read it as something else:
The obligation is fixed. The date is fixed. The instructions don’t exist yet.
That’s not a reprieve. That’s a gap with a deadline attached.
And here’s the trap. If you wait for the standard, you start building in 2027 and you’re building in the same year you’re supposed to be finished. I’ve watched enough enterprise timelines to know how that ends.
The work didn’t move at all
Forget the dates for a second. If you run production AI touching EU users, these questions are exactly what they were in June:
Which systems do you actually have, and what risk class is each one?
Most teams I talk to can’t produce this list. Not “can’t produce it quickly.” Can’t produce it. That finding is the finding.
What did each system log, and would that log survive someone who doesn’t trust you?
An application log, written by the system being audited, with no integrity guarantee, is not evidence. It’s an assertion in monospace.
Which model, which version, which data, which prompt produced which output?
For one API call this is annoying. For an agent pipeline with retrieval, tool calls and handoffs, most teams genuinely cannot reconstruct it afterwards. I’ve tried on real systems. It breaks fast.
Where did a human actually intervene, and what proves it?
A policy saying a human could intervene is not evidence that one did.
That’s roughly twelve months of architecture for a mid-sized estate.
It’s not a Q3 2027 procurement line item.
And it isn’t only Europe
This is where waiting gets genuinely expensive, because the pressure is now coming from three directions at once.
Saudi Arabia declared 2026 the Year of AI. SDAIA’s AI Adoption Framework has become a procurement filter in practice — model accountability, transparency, human oversight, evidenced. They published a National AI Risk Management Framework in April. SDAIA itself has held ISO/IEC 42001 certification since July 2024, one of the first government agencies in the world to get it.
UAE Central Bank issued responsible AI and ML guidance this year. Qatar Central Bank has its own AI guideline. ADGM has AI provisions in its rulebook.
Different regulators. Different languages. Same demand underneath: produce the evidence.
Build the architecture once and it serves all three.
Build it three times because three procurement teams asked in three different quarters, and you’ve burned a year.
The part I actually work on
I don’t find compliance interesting. I find the technical problem underneath it interesting.
Not whether to log. What a log needs to contain when the thing you’re proving passed through several hands.
One model answering one question — provenance is bookkeeping.
But when a retrieval step feeds a summariser, which feeds a decision agent, which calls two tools and hands off to a second agent, and the final answer is wrong — which link failed?
Standard observability won’t tell you. It tells you what ran, in what order, how long it took. It doesn’t tell you whose reliability you should have discounted, or how much confidence you’re entitled to given the weakest thing that touched the output.
That’s not a logging problem. That’s a transmission problem.
Which is a problem somebody already solved.
Old science, new problem
Classical hadith scholarship spent roughly twelve centuries on a structurally identical question:
How do you assess a report that reached you through a chain of narrators, none of whom you can interrogate directly?
The system they built — isnād–rijāl — is formal, not vibes. Grade every narrator against documented criteria. Evaluate the chain. Let the weakest link constrain how much confidence you assign to the claim.
Swap “narrator” for “model, retriever, scraper or tool.”
The machinery transfers almost intact.
I formalised it for multi-agent AI, published the paper (arXiv:2607.24117), and open-sourced the implementation (pip install isnad).
Then I tested whether the analogy actually holds
Because an elegant analogy is worth nothing on its own, nobody had checked whether the grading rule survives contact with data.
So v2.6.0 ships ISNAD-Bench: the library’s weakest-link rule run against 577,024 real hadith transmission chains, scored against the verdicts classical scholars actually reached on the same chains.
Results:
- Cohen’s κ = 0.871 (strict configuration)
- κ = 0.761 (lenient)
- κ = 0.331 between the human scholars themselves, on the same material
Two things about that third number, because it’s the interesting one and it’s easy to abuse.
The honest caveat first: inter-scholar κ and framework-vs-consensus κ are not measuring the same thing. Scholars disagreeing with each other pairwise is a harder test than an algorithm matching an aggregated verdict. I’m not claiming the rule beats scholars. Anyone who tells you a number like that is selling something.
What I am claiming is narrower: this domain has real, irreducible expert disagreement — and a mechanical weakest-link rule still reproduces the settled consensus at high agreement.
The method survives half a million real cases.
That’s what I actually wanted to know. And it’s the result that would have killed the project if it had come out the other way.
The benchmark changed my mind about something
This is the part I didn’t expect.
The bench made me flip the library’s default.
Strict is now the default. Ungraded narrators cap at ḍaʿīf — weak. lenient_unknown=True is opt-in.
Before, an unknown contributor got the benefit of the doubt. That felt generous and reasonable and it was wrong.
An unknown contributor should cost you confidence, not leave it untouched.
I had an intuition. The numbers disagreed with my intuition. The numbers won. That’s the whole reason to build a benchmark.
What this does not do
I’d rather say this myself than have someone say it to me in the comments.
ISNAD produces evidence artifacts. It does not make anyone compliant with anything. Not the AI Act, not ISO 42001, not SDAIA’s frameworks. No tool does. Conformity gets assessed against standards that, as covered above, are still being written.
Chain independence is still an open problem. The weakest-link rule assumes contributors are meaningfully distinct. In agent systems sharing a base model or a retrieval corpus, that assumption strains. It’s tracked publicly in the repo. I haven’t solved it.
The benchmark validates the grading rule against a historical corpus. It doesn’t prove the rule transfers cleanly to model-generated content, where “narrator reliability” is a much younger and slipperier idea. That’s the next piece of work.
If you’re going to use something in an audit trail, you should know exactly where it’s weak. So there it is.
What I’d do in the next 90 days
Not in 2027.
1. Inventory. Every AI system in production, with a provisional risk class. If you can’t do this in a week, you just learned something important.
2. Check your Article 50 exposure. It’s already live. If you ship generative content, the 2 December marking deadline is fourteen weeks out.
3. Pick one system and try to reconstruct a single decision end to end. Model, version, data, prompt, human review. Whatever breaks is your architecture gap. It’ll break faster than you think.
4. Design the evidence layer once, for all three jurisdictions. The EU, SDAIA and the Gulf central banks are asking for the same artifact in different dialects.
5. Don’t wait for the harmonised standard. Build something defensible now, map it to the standard when it lands.
Demos are easy. Production is where things break. Compliance theatre is easy too. Evidence is where it breaks.
Old science. New problem. Same question underneath both:
Can you prove where this came from?
Paper: arXiv:2607.24117 · Code: pip install isnad · Benchmark: ISNAD-Bench, v2.6.0
Regulatory dates verified 26 August 2026. This is a technical article, not legal advice.
If you’re running production AI and can’t answer “what did it do and how do you know” — that’s a conversation for this quarter, not 2027. I’m easy to find.