Home / Insights / What to Measure in an AI Pilot
Guide

What to Measure in an AI Pilot

Summarize with AI Prompt copied — paste it into the chat

Most AI pilots are judged on a demo. Someone shows the tool working, the room agrees it looks impressive, and the decision to continue gets made on a feeling. Three months later nobody can say whether it helped, because nothing was measured before it started.

Measurement is the cheapest part of a pilot and the part most often skipped. Here is what to put in place, and when.

Before anything is built: the baseline

You cannot show improvement against a number you never recorded. The baseline has to be captured before the pilot starts, because afterwards the process has changed and the original figure is unrecoverable.

Keep it simple and specific to the process being piloted. If the pilot drafts first responses to warranty emails, the baseline is how long a first response currently takes and how many drafts get materially rewritten. Not "customer satisfaction" — a number that moves for a dozen reasons is a number the pilot cannot claim credit for.

  • Measure for at least two normal weeks. One week catches whatever was unusual about that week.
  • Record who measured it and how, so the comparison at the end is like-for-like.
  • Write down the number in a shared place before the build starts. Baselines remembered after the fact drift towards whatever makes the result look good.

The pass mark, written before you see results

Decide in advance what result would justify moving to production. This is the single highest-leverage thing on this list, and it costs nothing.

The reason is psychological rather than technical. After a pilot, a good outcome and a mediocre one look remarkably similar in a meeting — both are presented by people who worked hard, both come with caveats, and there is always a case for one more iteration. A number agreed in advance is the only thing that reliably distinguishes them.

A usable pass mark is specific and falsifiable: "the drafted response is sent without material edits in at least 60% of cases, measured over two weeks". Not "the team finds it useful".

Counter-metrics: what must not get worse

Almost any process can be made faster if you stop caring about something else. A pilot that improves handling time while quietly increasing error rates is not a success, and it will be discovered in production rather than in the pilot.

So pair every primary metric with one thing that is not allowed to degrade. Speed with accuracy. Volume with quality. Cost with rework. One counter-metric is usually enough; the point is to make the trade-off visible rather than to build a scorecard.

Who owns the number

A metric with no owner does not survive the pilot. If nobody currently reports the figure, nobody will notice it improve, and nobody will make the case for production when the pilot ends.

This is a selection criterion as much as a measurement one. If you cannot name the person who owns the number a pilot would move, that is good evidence you have chosen the wrong process — not that you need a better dashboard.

What not to measure

  • Model accuracy on a benchmark. It tells you about the benchmark, not your inputs.
  • Usage. People will use a new tool during a pilot because it is new and they are being watched.
  • Anything that takes longer to collect than the pilot runs. If instrumenting the measurement is a project, measure something simpler.
  • Satisfaction scores as the primary metric. Useful context, too noisy to decide on.

Bringing it together

A well-instrumented pilot needs four things written down before the build: the baseline, the pass mark, one counter-metric, and the name of the person who owns the number. That is perhaps an hour of work, and it is the difference between a pilot that concludes and a pilot that merely ends.

Our €2,500 audit produces exactly these four before any code is written, which is also why it sometimes concludes that the process is not worth piloting at all.

Frequently asked questions

How long should the baseline period be?

At least two normal weeks, and longer if the process is seasonal or has a monthly cycle. One week tends to capture whatever was unusual about that week rather than the process itself. If the work has a clear month-end peak, either measure a full month or deliberately exclude the peak from both the baseline and the comparison — but decide which before you start, not after.

What if we have no baseline data at all?

That is common and it is fixable, but it costs two weeks at the front of the pilot rather than being skipped. Measure manually if you have to — a tally kept by the team for ten working days is a perfectly respectable baseline. What does not work is reconstructing the number afterwards from memory or from a system that was not recording it, because the reconstruction will unconsciously flatter the result.

Should the pass mark be a percentage or an absolute number?

Whichever the owner of the number already uses. If the team reports hours saved, set the pass mark in hours; if it reports a rate, use a rate. Converting into a unit nobody uses day to day is a small friction that reliably turns into an argument at the decision meeting.
Our AI services Hire an AI consultant AI automation AI agents AI implementation Pricing

Want any of this applied to your business?

We turn these concepts into working tools — grounded, safe and measurable. Start with a free consultation.

Book a free consultation →