Building software became affordable, so trying things became affordable.
An AI pilot is a limited live trial of an AI solution with real users, usually one team, department or customer group. It tests how the system performs in practice, surfaces adoption issues, and gathers feedback before a wider rollout. Piloting reduces risk: problems are found on a small scale and fixed before the solution reaches the whole organisation.
An AI pilot is a small, time-boxed build that runs on your real data, with your real users, to answer a question you cannot answer on paper: does this work here, and is it worth doing at scale?
Pilots exist because of an economic shift, not a fashion. Building software has become dramatically less expensive and faster. That did not just reduce the cost of projects, it changed which strategy is rational. When a build cost €250,000 and took a year, you had to be certain before you started, and certainty was expensive to buy. Now that the same capability costs a fraction of that, the less expensive move is to stop trying to be certain and start finding out.
For thirty years the sensible approach to software was to specify heavily up front, because changing your mind mid-build was ruinous. That logic held while building was the expensive part. It no longer is. Modern tooling, foundation models and managed infrastructure have collapsed the cost of a first working version, and the expensive part is now knowing what to build. When the answer costs more than the experiment, run the experiment. That is the whole argument for piloting, and it is why the practice arrived with AI rather than before it.
Information, not software. A pilot that ends with a working tool but no clearer decision has failed; a pilot that ends with no tool but a firm, evidenced "no" has succeeded and saved the production budget. Three things are usually worth more than the prototype: whether your data supports the use case (it frequently does not, and this is the most common killer), whether the people whose work changes will actually use it, and what the real error rate is on your inputs rather than a benchmark.
1. One process, named. Not "AI for customer service" but "the first-response draft for warranty emails". 2. A number someone already owns. If nobody currently reports the metric, nobody will notice it improve. 3. A pass mark written down first. Decide what result would justify production, before you see any results. 4. A real user, not a stakeholder. The person who does the work daily, not the person who approved the budget. 5. A stop date. A pilot without an end date becomes a permanent side system that nobody maintains.
Low-cost experiments create a failure the expensive era did not have. The old risk was betting €250,000 on the wrong thing. The new risk is running eight pilots and shipping none of them. Each one looked promising, none had a written pass mark, so none could be declared finished or dead. The discipline that matters now is not the entry criteria, it is the exit criteria. A pilot needs a defined way to fail as much as it needs a way to succeed, and someone with the authority to call it either way on the stop date.
You will have seen the number: 95% of enterprise GenAI pilots deliver no measurable P&L impact. It comes from The GenAI Divide: State of AI in Business 2025, published by MIT's NANDA initiative, which drew on 52 executive interviews, a survey of 153 leaders and a review of roughly 300 public deployments. Gartner published a second figure the same year: more than 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.
Both numbers are worth reading properly rather than repeating. The MIT report defined success narrowly and deliberately: deployment beyond the pilot phase, with measurable KPIs, and ROI assessed six months after the pilot ended. On that definition, efficiency gains, lower churn and faster pipeline all count as nothing. The report also describes its own interview findings as directionally accurate rather than company-reported. So 95% is not 95% of pilots collapsing; it is 95% failing a strict test that most pilots were never set up to sit.
Which is the useful part. That strict test is exactly the bar to design for. If a pilot cannot name the KPI it moves, the person who owns that KPI, and the date six months out when someone checks, it will land in the 95% by construction, whatever the model does. Both findings point at the same cause, and it is not model quality: it is integration into a real workflow, and deciding in advance what would count as proof.
At Crux Digits: a €2,500 audit first, to establish whether the process is worth piloting at all: roughly one to two weeks, and it regularly ends with "a €4,000 script does this, do not build AI". Then a €20,000 Production-ready MVP running on your data for 4–6 weeks, and production from €50,000 if the pass mark is met. Fixed prices, because a pilot whose cost is open-ended is not time-boxed in any meaningful sense.
The pilot in six weeks
Building software became affordable. That makes finding out less expensive than being certain: provided you agree up front how it ends.
When a build cost €250,000 you had to be certain before starting. Now a first working version costs a fraction of that, finding out is less expensive than being sure.
Not “AI for customer service” but “the first-response draft for warranty emails”. A pilot spanning two processes answers neither.
If nobody reports the metric today, nobody will notice it improve, and nobody will push it into production.
Decide what result would justify production before you see any results. Afterwards, a good outcome and a mediocre one look identical.
The person who does the work daily, not the one who approved the budget. A formal role, not a demo at the end.
Four to six weeks, fixed at the start. Without an end date a pilot becomes a permanent side system nobody maintains.
On the stop date the owner decides. Running eight pilots and shipping none is the new failure mode, not betting on the wrong thing.
Four to six weeks of build, on a stop date fixed at the start. Long enough to hit real data problems, short enough that the business context has not moved by the time you report.
A proof of concept asks "can this work at all?" and can be answered in a lab, on sample data. A pilot asks "does this work here, with our data and our people?" and can only be answered in the business.
The person who owns the number it is trying to move, not IT, and not an innovation function, unless one of those happens to own that number.
Want this applied in your business? See how we take it to production:
We build this AI in production, at fixed prices, with one named expert. Start with a free consultation.
Book a free consultation →