Case studies · Benchmarks

The systems, with the numbers.

Client case studies will be published here with their owners’ consent — named, measured, method stated. Until then: three systems from our own operation, written to the same standard we’ll hold client work to.

Case 01 — Content systems

The newsroom pipeline.

The newsroom pipeline — a daily publication run by agents

What it is: a scheduled pipeline that captures AI-industry sources, scores and shortlists stories, drafts analysis, and publishes to our research site — every weekday, with a human gate before publish.

The numbers: 76 articles published in 43 days (8 July – 19 August 2026), roughly 1.8 per day. Each article carries a provenance chain — source ID, run ID, and a cryptographic hash before export and after publication — so any piece can be traced back to its inputs, byte for byte.

Read the newsletter →

Case 02 — Research at scale

The qualification engine.

The qualification engine — a pipeline with receipts

What it is: a governed pipeline that discovers businesses in a territory, captures their public web presence, and scores each one — with the evidence retained for every record.

The numbers: 4,703 UK businesses assessed across nine towns and 445 categories. Of those: 3,203 with a verified owned website, 2,548 with clean measured website scores, 4,412 with reputation scores, and 3,478 distinct evidence files retained by hash. Every stage reconciles — records are never silently dropped, and every score can be replayed from its evidence.

Why it matters to you: this is what “assessed, with receipts” looks like at a scale no one does by hand — the same discipline we bring to a client’s data.

Case 03 — Model benchmarks

We test AI before we trust it.

Benchmark testing — we test AI before we trust it

What it is: before a model runs our work — or yours — it runs our benchmarks: identical tasks, fresh sessions, scored from the actual execution trace against answers we establish first. A model’s own report is never taken as evidence.

What three rounds found (August 2026): one model failed the same task twice without ever reporting it was blocked. One model family skipped a required declared-planning step in eight runs out of eight. Two models invented completion timestamps that predated their own start times, in four runs out of four. One broke a read-only boundary during an investigation exercise.

Why it matters to you: these are failure classes a sales demo never shows. Knowing which model behaves under which conditions is the difference between automation you can leave running and automation you babysit. We publish behaviour classes, not vendor scoreboards — the tests are ours, on our tasks, on dated versions.

Ask how any of this was measured.

Every number on this page has a method, a date, and an evidence trail behind it.