Topic hub

Verifiable AI Systems

A system is verifiable when a material result can be connected to a source, data version, method, test, and human decision. This is not a certificate of infallibility; it is a property of architecture and process.

01

Position

How I use this term

A system is verifiable when a material result can be connected to a source, data version, method, test, and human decision. This is not a certificate of infallibility; it is a property of architecture and process.

A model should be treated as one component of an evidence path. Admissible sources and quality criteria come first; retrieval, models, and tools follow.

The important question is not whether a model can produce a correct answer in a demonstration. It is whether a team can reconstruct the conditions that produced the result, detect a regression after change, and name the person accountable for the decision. Verifiability therefore covers data, execution, testing, and a procedure for challenging an outcome.

I use this hub as a map of practice, not an automatically generated tag collection. The starting point is a concrete user task, admissible material, and a failure that must not pass unnoticed. Model, retrieval, and tool choices come later. Every account should separate observation, assumption, interpretation, and human decision. This lets a reader inspect not only a proposed solution but also the conditions under which it stops being dependable.

Evidence has several layers here: source and licence, data structure, execution version, evaluation scenario, and accountability path. A project should lead to a public artifact, while an article should connect to a project or research object. In this cluster those checkpoints include NormaLab, BY Maps. Status remains explicit, and limitations stay in the main account rather than appearing as a disclaimer after a demonstration.

A useful reading path begins with Retrieval crash test: when similar means wrong, Fundamental LLM problems in legal systems and continues into the related case studies. The sequence is not a sales funnel; it shortens the route from a concept to an inspectable example. A performance claim needs a benchmark. A source-quality claim needs a corpus and negative cases. A claim about professional decisions must identify the oversight point and a practical way to challenge the result.

The boundary of the field matters as much as its definition. Not every automation needs an LLM, not every artifact can be made public, and a prototype is not evidence of production readiness. The material therefore carries a version, date, status, and scope note. The readiness audit below translates those principles into a visitor’s own workflow by checking sources, baseline, evaluation, accountability, and a safe change procedure before model integration begins.

An update to this hub should begin with a change in evidence, not a desire to add another label. A new model, source, or procedure requires a check of which conclusions still hold, which localisations have become stale, and whether every link still resolves to the same artifact version. In practice that means a small change register, a review date, and a named trigger for rerunning the test. This discipline lets the topic develop without concealing earlier errors and keeps a durable method separate from a one-off experiment. A reader can then assess not only the current result but also how it changed after criticism or a material change in the data.

02

Projects

Where the method is used

Pilot / Legal AI and Computational Law

NormaLab

Conventional legal search retrieves similar documents. NormaLab asks a different question: how did a specific rule operate across a body of judgments, where is practice stable, where does it diverge, and what supports each conclusion?

The public method exposes the path from question and corpus to citation, report, and evaluation verdict—even when the work is returned for revision.

My role

Data engineering · Agent systems · Reproducibility · Citation engine

Evidence

live demo · article

CASE / NORMALABcase-v1.1Updated

Research / Open and Applied Research

BY Maps

BY Maps is a public research material on demographic change, combining a map, research questions, and conditional scenarios. It keeps each result close to its source, transformation, and limitation.

The public material connects a geographic view with links to sources, method, and limitations. Conclusions require checking against versioned data.

My role

Research framing · Data pipeline · Modelling · Visualisation

Evidence

live demo · repository

CASE / BY-MAPSresearch-v1.1Updated
03

Start here

Cornerstone material

04

Reading

All writing in this area

Verifiable AI Readiness Audit

Apply the method to your workflow

Check sources, data, evaluation, oversight, and security before selecting a model.

Start the diagnostic