Evidence

Every claim on this site links to where it came from.

No adjectives stand in for proof. Here are the artifacts you can open yourself.

Open any of these

Hadith benchmark · κ = 0.87

The weakest-link rule run against 575,060 scholar-graded chains. κ = 0.8714 (3-way; 1 = perfect, 0 = chance) — agreement with a rule-based convention (Ibn Hajar's 12 tiers), not ground truth. 5-way 0.8667, lenient 0.761, narrator-grade κ = 0.33, shuffled control ≈ 0 (-0.0066). 577,024 is the raw corpus of sanads — the benchmark N is 575,060.

Results in the repo →

RAGTruth transfer · κ = 0.4345

1,800 responses from 6 models: 96.0% hallucination recall at 60.4% precision (recall ≠ precision), 70.3% accuracy against a 55.8% baseline. 96.8% parsed; the 3.2% that did not parse are scored fail-closed, not excluded.

Read case #0 →

What evidence actually means

1

AI system inventory — every model, agent, and pipeline in use, versioned and located.

2

Risk classification — which systems are high-risk, under which framework, and why.

3

Decision and event logging — what the system did, when, and on what input.

4

Provenance and lineage — where each claim came from, through whom, and whether the chain holds.

5

Human-oversight evidence — who reviewed what, and the trail that proves it.

6

Technical documentation pack — the artifacts an auditor can actually open.

Deadlines are already set

European Union

Annex III high-risk obligations 2 December 2027; Annex I 2 August 2028 (moved by the Digital Omnibus, in force 27 July 2026).

Art. 50 transparency duties for deployers have applied since 2 August 2026.

Art. 12: record-keeping / logging capability · Art. 19: logs kept at least six months · Art. 18: documentation kept 10 years.

Source: EUR-Lex · Regulation (EU) 2024/1689
Saudi Arabia

SDAIA published its AI Adoption Framework 12 September 2024 to encourage government entities, and a National AI Risk Management Framework on 14 July 2026.

Source: SDAIA / Saudi Press Agency
Gulf financial services

UAE Central Bank: Guidance Note on the Consumer Protection and Responsible Adoption and Use of Artificial Intelligence and Machine Learning by Licensed Financial Institutions in the U.A.E. (issued 11 February 2026).

Source: Central Bank of the UAE rulebook

The delay is not relief. Evidence architecture takes months to build, and the technical standards are still being drafted. That gap is the work.

Open source, already running

ISNAD

What happened, through whom, and which link was weakest. Listed in LangChain's middleware integrations.

ISNAD →

knowledge-ci

Has answer quality degraded — and fails your build.

knowledge-ci →

RAGLint

Is this data fit for retrieval.

RAGLint →

Agent-Surface

Can an agent actually use this interface.

Agent-Surface →

The way in

Free 30-minute call

For qualified teams.

We check the four-question filter and look at your pipeline. No deck, no pressure.

Book the call →

Provenance Assessment

two weeks · one pipeline · $5,000 (first 2) · €7,500 after

A technical readiness review with a written report, signed evidence, and a coverage map.

See the process →

Annual licence & ongoing

For teams that want it to keep running.

ISNAD in your pipeline, the evidence kept current, and me on call for the hard questions.

Let's talk →

A technical readiness review — not an Article 43 conformity assessment and not a notified body.

Trusted by people I've built with

"Ali is an exceptionally talented software engineer with deep expertise in end-to-end system design and a passion for advancing AI and ML technologies. His ability to design and build AI systems is outstanding."

Noah Marra

Sr. Machine Learning Engineer, AMD · Former teammate

"Ali Raja is an exceptional and talented data scientist that possesses great entrepreneurial and technical abilities. Working with him is always a pleasure."

Ali Almussa

Entrepreneur & Project Management Consultant