Most AI pilots are judged on a demo. Someone shows the tool working, the room agrees it looks impressive, and the decision to continue gets made on a feeling. Three months later nobody can say whether it helped, because nothing was measured before it started.
Measurement is the cheapest part of a pilot and the part most often skipped. Here is what to put in place, and when.
Before anything is built: the baseline
You cannot show improvement against a number you never recorded. The baseline has to be captured before the pilot starts, because afterwards the process has changed and the original figure is unrecoverable.
Keep it simple and specific to the process being piloted. If the pilot drafts first responses to warranty emails, the baseline is how long a first response currently takes and how many drafts get materially rewritten. Not "customer satisfaction" — a number that moves for a dozen reasons is a number the pilot cannot claim credit for.
- Measure for at least two normal weeks. One week catches whatever was unusual about that week.
- Record who measured it and how, so the comparison at the end is like-for-like.
- Write down the number in a shared place before the build starts. Baselines remembered after the fact drift towards whatever makes the result look good.
The pass mark, written before you see results
Decide in advance what result would justify moving to production. This is the single highest-leverage thing on this list, and it costs nothing.
The reason is psychological rather than technical. After a pilot, a good outcome and a mediocre one look remarkably similar in a meeting — both are presented by people who worked hard, both come with caveats, and there is always a case for one more iteration. A number agreed in advance is the only thing that reliably distinguishes them.
A usable pass mark is specific and falsifiable: "the drafted response is sent without material edits in at least 60% of cases, measured over two weeks". Not "the team finds it useful".
Counter-metrics: what must not get worse
Almost any process can be made faster if you stop caring about something else. A pilot that improves handling time while quietly increasing error rates is not a success, and it will be discovered in production rather than in the pilot.
So pair every primary metric with one thing that is not allowed to degrade. Speed with accuracy. Volume with quality. Cost with rework. One counter-metric is usually enough; the point is to make the trade-off visible rather than to build a scorecard.
Who owns the number
A metric with no owner does not survive the pilot. If nobody currently reports the figure, nobody will notice it improve, and nobody will make the case for production when the pilot ends.
This is a selection criterion as much as a measurement one. If you cannot name the person who owns the number a pilot would move, that is good evidence you have chosen the wrong process — not that you need a better dashboard.
What not to measure
- Model accuracy on a benchmark. It tells you about the benchmark, not your inputs.
- Usage. People will use a new tool during a pilot because it is new and they are being watched.
- Anything that takes longer to collect than the pilot runs. If instrumenting the measurement is a project, measure something simpler.
- Satisfaction scores as the primary metric. Useful context, too noisy to decide on.
Bringing it together
A well-instrumented pilot needs four things written down before the build: the baseline, the pass mark, one counter-metric, and the name of the person who owns the number. That is perhaps an hour of work, and it is the difference between a pilot that concludes and a pilot that merely ends.
Our €2,500 audit produces exactly these four before any code is written, which is also why it sometimes concludes that the process is not worth piloting at all.