AI implementation for real workflows: from readiness to a measurable pilot
AI implementation for high-stakes workflows: readiness assessment, source-grounded RAG and agent design, evaluation, and a bounded measurable pilot.
01
AI problem diagnostic
When it is useful
The team has an idea, requirements, or a prototype, but no shared answer to which decision should improve, which data is usable, which errors are unacceptable, or how value will be measured.
What we do
We examine the real workflow, participants, sources, constraints, baseline, and cost of error. Model tasks are separated from organisational and human work.
What the team receives
A workflow and data map; a register of material risks; quality criteria; a bounded pilot option—or a reasoned decision not to use AI.
A general-purpose chat interface is insufficient: results must use owned sources, respect versions, preserve citations, and pass defined checks.
What we do
We design the corpus, data model, retrieval, tools, agent roles, stopping points, and human decisions. We specify how the system will detect a wrong document, incomplete context, or superseded version.
What the team receives
An architecture specification; data-access model; result contract; quality-evaluation plan; pilot boundaries and readiness criteria.
A large corpus of legislation, judgments, case files, or professional documents must be analysed without turning textual similarity into a legal conclusion.
What we do
We define the research question and unit of analysis, assemble and version the corpus, and design classification, citation checks, counterexample search, and legal review.
What the team receives
A corpus plan; processing methodology; source and citation workflow; structure for a verifiable report; and a pilot or research-programme plan.
An organisation is considering a broad AI initiative but first needs to test its central hypothesis on a real workflow.
What we do
We bound the work to one observable outcome, assemble a control set, build the smallest working prototype, and test both typical performance and critical failure modes.
What the team receives
A working prototype; baseline and test results; a register of known limitations; material for a go / reshape / stop decision; and a handover plan for the internal team.
I provide practical reviews and project mentoring in RAG, agent workspaces, AI-native development, and research pipelines. Participants work with their own workflow, repository, and verifiable outcome—not a training chatbot.
I do not promise automation before understanding the workflow.
The work is unlikely to help if the requirement is unchecked mass content, a public demo without user access, accuracy promised in advance, or a deployment that nobody inside the organisation is prepared to maintain. In those cases, the problem definition needs to change first.