An AI exit strategy is not a clause in the contract. It is three things your IT organisation keeps for itself: the evaluation set that defines what a correct answer is, the data contracts that say what the system may read, and the registry of every prompt and model version in production. Whoever holds those three can change AI supplier in weeks. Everything else, starting with the build, you can outsource, and should.
This essay is for the IT manager, IT director or CIO at an organisation of 250 to 5,000 people who already has an ICT partner and is now bringing in a specialist AI supplier next to it. I run one of those specialists, so read what follows knowing that it argues against the easiest way for a firm like mine to keep a client. I think it is the right line anyway, and I would rather draw it for you than have you discover it during a supplier change.
Why is "keep the strategy, outsource the build" the wrong line?
The advice most IT departments receive sounds sensible. Keep the strategy in house, let the specialist do the building. It fails for a plain reason: a strategy is a document, and nothing in production ever checks itself against a document.
What actually decides whether you can leave a supplier is not who wrote the roadmap. It is who can prove, on a Tuesday morning, that the system still does what it did last month. In a conventional application that proof lives in the specification and the test suite, and your ICT partner has handled both for years. In an AI system the specification is fuzzy by nature, the output is a judgement rather than a calculation, and the only honest definition of "working" is a set of real cases with the answers you expect. If that set lives with the supplier, then so does the definition of your system.
I see the same pattern in most first conversations with a mid-sized IT function. The organisation owns the source code on paper, the contract has an exit clause, and the architecture diagram is in the CMDB. Then I ask who could tell whether a replacement supplier's version is as good as the current one, and the room goes quiet. Nobody can, because the evidence of "as good" was never theirs.
What did OpenAI's June 2026 deprecations say about where your evals live?
On 3 June 2026 OpenAI told developers that its Evals platform is being deprecated. According to the OpenAI deprecations page, existing evals become read-only on 31 October 2026 and the Evals dashboard and API are scheduled to shut down on 30 November 2026. The same day it announced that reusable prompt objects and the v1/prompts API will shut down on 30 November, with the advice to move prompt content into your own application code, and that Agent Builder will close on the same date.
I do not read that as a scandal. Platforms retire products, OpenAI published dates months ahead and pointed to a migration path. The point is narrower and more useful. The two artefacts I would tell any IT function to hold itself, the evaluation set and the prompt registry, are exactly the two a major model provider announced in June it will stop keeping for its customers. A team that let a vendor dashboard be the only home of its test cases and graders now has a deadline it did not choose. A team whose evaluation set sat in its own repository reads the notice, shrugs and carries on.
Notice also what the advice was: move the prompts into your own code. The provider is telling you where they belonged all along.
What should IT keep in-house with an AI supplier?
Three artefacts. None of them needs a data scientist to own. All of them need a named person inside your organisation.
1. The evaluation set
This is a versioned collection of real inputs from your own process, each with the answer you would accept and a rule for grading it. For a service-desk assistant that means actual tickets with the correct routing. For a document classifier it means actual documents with the correct label, including the awkward ones. The business owner of the process decides what counts as correct. IT decides where the set lives, who may change it and how every run is recorded.
The evaluation set is also what turns a wrong answer into something your service desk can handle. When a user reports that the assistant got something wrong, the case goes into the set, and from then on every model or prompt change is tested against it. I wrote about why a wrong answer falls outside every existing definition of an incident in AI incident ownership. The evaluation set is the instrument that closes that gap, and it only closes it if you hold it.
2. The data contracts
An AI system reads from systems your ICT partner already runs: the ERP, the document store, the CRM, the ticketing tool. A data contract states which fields the AI system may read, from which source, how fresh they must be, and who must be told before the schema changes. It is a short document plus a check that fails loudly when the contract is broken.
This is the seam between your two suppliers, and a seam belongs to the party that sits on both sides of it. The AI supplier cannot own it, because it does not control the source systems. The ICT partner will not own it, because it does not know what the model depends on. You are the only party that sees both.
3. The prompt and model registry
Every prompt in production, every model and model version, every retrieval setting and parameter that changes behaviour, each with a date, an author and a reason. It is the change record for the part of an AI system that changes without a release. When a provider upgrades a model underneath you or a supplier tunes a prompt on a Friday, the registry is how you know, and the evaluation run attached to the entry is how you know whether it mattered.

A registry does not need a product. A repository with one file per prompt and a log of evaluation runs is enough for most organisations to begin with. What matters is that your change process can see it.
Is owning the source code not enough?
Owning the code matters, and it should be in the contract. I argued in AI consultant lock-in that a supplier should hand over code, documentation and data ownership as a matter of course. But in an AI system the code is the least specific part. A competent new supplier can rebuild the integration in a few weeks. What they cannot rebuild is two years of accumulated knowledge about which answers your organisation accepts, which fields broke the model last spring and why the prompt says what it says. That knowledge sits in the three artefacts or it sits in someone's head at the old supplier.
What can a specialist AI supplier own without hurting you?
Almost everything else. The build and the integration. Model selection and the trade-offs behind it. Retrieval design, tuning and the day-to-day improvement work. Monitoring tooling, and even running the system if you want that, provided every change is tested against your evaluation set and written to your registry.
This is where I disagree with the instinct to insource AI capability wholesale. Most organisations in this size band do not need an internal machine-learning team to be safe. Whether an in-house team makes sense is a separate commercial question, which we set out in AI agency versus in-house team. What they need is the ability to judge the work, which is a smaller and very different thing. It is closer to the role of a good principal on a construction project than to that of a builder.
The same holds for your ICT partner. It keeps the estate, the identities, the network and the change process it already runs. Nothing in this line asks it to become an AI firm, and nothing asks the AI supplier to become an ICT partner.
What did the Dutch state learn about keeping judgement in-house?
The Netherlands has already run this experiment at scale, in public. On 15 October 2014 the temporary parliamentary committee on government ICT, the Elias committee, delivered its final report, Naar grip op ICT. It concluded that central government did not have control over projects with a major ICT component. Its recommendations covered, among other things, the ICT knowledge inside central government itself, and procurement and contract management. Its best-known recommendation became the Bureau ICT-toetsing, set up in July 2015 to assess central-government projects above five million euros before they start. Its successor, the Adviescollege ICT-toetsing, was put on a permanent statutory footing on 1 July 2024.
The lesson I take from that history is modest and, I think, durable. The state never stopped outsourcing the building of systems, and it was never going to. What it had lost, and spent a decade rebuilding, was the capacity to judge on its own behalf whether what it was buying would work. It built an assessor, not a software house.
The analogy only travels so far. Your AI assistant is not a multi-year public programme, and the Elias committee wrote about ICT in general rather than AI. It is also a Dutch story: Flemish readers have their own public-sector history and should not read it as theirs. But the shape is the same one I see in private IT functions today. The building is outsourced, correctly. The judgement goes with it, by accident.
How do you tell whether you could switch AI supplier?
One test, and you can run it this week. Could you hand your evaluation set, your data contracts and your registry to a different supplier on a Monday, and have them show by Friday that their version performs at least as well as the current one, on your cases, by your grading rule?
If the answer is yes, you have an exit strategy, whatever the contract says. If the answer is no, the contract clause is decorative. Three questions to put to your current AI supplier make the gap visible quickly:
- Where does the evaluation set for our system live, and could we run it without you?
- Which data fields does the system depend on, and what happens when one changes?
- Show us every prompt and model change in the last ninety days, and the evaluation result for each.
A good supplier answers all three without hesitation and is glad you asked. If one resists handing over the evaluation set, it is telling you where it expects its margin to come from. The same questions apply if you ever move from a managed agent platform to your own stack, which is the technical version of this problem covered in managed agent APIs.
Why would an AI supplier argue against its own lock-in?
Because the alternative is a worse business. A supplier that keeps clients by holding their evaluation evidence has to be at least a little afraid of every conversation about quality. A supplier that hands it over has to keep earning the work. I prefer the second position, and it is the one we take: at the end of an implementation the client receives the code, the models and the evaluation set, which is how our AI implementation partner work is scoped.
There is a cost to your side, and I should name it. Keeping these three artefacts alive is a role, not a department, but it is a real role. Someone has to add the failed cases, review the registry and chase the data contract when a source system changes. In most organisations of this size that sits naturally with an information manager or a functional application manager who already understands the process. If nobody can be named, that is the finding, and it is better to learn it before go-live than after a supplier change.
Where should an IT function start this quarter?
Start with the AI system you depend on most, not the newest one. Ask your supplier for the evaluation set and get a copy into a repository you control. If there is no set, that is your first piece of work, and the business owner of the process has to be in the room for it. Then write the data contract for the two or three source fields the system cannot live without, and put a check on them. Finally, open the registry: one entry per prompt and model in production, and a rule that nothing changes without an entry and an evaluation run.
None of that requires a new tool, a new team or a new supplier. It requires deciding, before anyone asks, which part of your AI estate is yours. For the broader picture of how mid-sized organisations are organising applied AI, see applied AI for mid-sized businesses.
Written by Tom Joseph, founder of Crux Digits. Checked against the OpenAI deprecations page and the AcICT site on 28 September 2026.
The AI audit & strategy: in one to two weeks, a ranked list of use cases with a return per opportunity, and a plan you own.
AI audit & strategy →


