In IT service management, AI change management is deciding which changes to an AI system go through your change process and which do not. Sort AI changes by what they can break, not by which file they touched. A new document in the knowledge base goes nowhere near the CAB. A one-line prompt edit that removes a refusal belongs there. And a behaviour change on the supplier's side is not a change you approve. It is one you detect.
This is written for the IT manager, change manager or CIO at an organisation of 250 to 5,000 people, where a change advisory board already meets every week and a specialist AI supplier is about to start putting work into production. By change management I mean ITIL change enablement: RFCs, change types, the CAB. Not the organisational kind about people and adoption, which is a different essay. I run one of those specialist suppliers. What follows is the classification we propose to clients, and I will say where it costs us something.
Why does an AI system change without a release?
ITIL change control grew up around an artefact. Something is built, tested, packaged and deployed; the RFC describes the artefact and the CAB judges the risk of putting it into production. The whole apparatus rests on one assumption: if nothing was deployed, nothing changed.
An AI system breaks that assumption in four places.
- The prompt. A system prompt is text, it is often stored as configuration, and editing it changes what the system says to every user from the next request onwards. No build, no deployment, no version number unless someone gave it one.
- The knowledge base. Add a document to the collection a retrieval system draws from and its answers change, without a line of code being touched.
- The model. Moving to a newer model version is the change that looks most like a classic release, and it is the one that goes wrong most often when it is treated as routine.
- The supplier's serving layer. This one surprises people. Anthropic's own documentation states that a model ID is a pinned snapshot whose weights do not change, but that the serving infrastructure around it, such as the request router, safety classifiers and sampling logic, can, and that such updates occasionally produce minor differences in observable behaviour. That is a model supplier being candid, and it is the right sentence to read before you write an SLA. Production behaviour can shift on a day when nobody on either side of the contract deployed anything.
So the organisations I talk to tend to end up in one of two places. Either nothing AI-related reaches the CAB, because nothing was "deployed", and the system drifts under nobody's signature. Or everything does: the CAB spends forty minutes on the wording of a prompt, the next wording change goes live on a Friday afternoon without an RFC, and the process quietly loses the argument. Both fail. The second fails more slowly, which is why people think it works.
What did the Dutch beheer model assume that AI breaks?
Dutch IT organisations already have the vocabulary for this problem. They have just not pointed it at AI yet. From 1986 onwards Maarten Looijen, later a professor at TU Delft, set out the three-way split of beheer that most Dutch IT departments are still organised around: functioneel beheer on the business side, applicatiebeheer for the software, technisch beheer for the infrastructure. The split found its way into ITIL, ASL and BiSL, and it lives on in job titles such as functioneel beheerder.
The split rests on a boundary that has held for decades. Functioneel beheer changes how an application is used and set up; applicatiebeheer changes what it does. A functioneel beheerder can add a product code, adjust a workflow setting or rewrite a letter template without an RFC, because none of those touch the logic.
A prompt sits exactly on that boundary, and breaks it. It is edited like a letter template, by someone on the business side, often in a text field. It behaves like code. The sentence "never quote a delivery date" is a piece of logic, and deleting it changes what the system may promise a customer. In an AI system, the text is the logic.
A note for Flemish readers: Looijen is a Dutch reference. If your organisation does not use his terms, the boundary he describes, between configuring an application and changing its logic, is still the one every ITIL-run IT department draws.
How should you classify AI changes? Three tiers by what they can break
The rule I propose: a change's tier is set by the worst thing it can break, not by the size of the edit or the type of artefact. In practice it sorts out like this.
Tier 1: changes to what the system knows
Content changes: adding or updating documents in the knowledge base, correcting an FAQ answer, refreshing a product catalogue feed. These change which facts come back, not the rules for handling them. Worst case: a wrong or outdated answer on one topic, the same risk as a wrong page on the intranet.

Route: no RFC. The owner is the content owner in the business, functioneel beheer in Looijen's terms. The gate is an automatic run of the evaluation set after each index update, with a threshold. The record is the index version in the log, so an incident can be traced back to what the system knew that day.
Tier 2: a standard change for how the system behaves within its limits
Prompt edits to tone, format or structure; retrieval settings such as how many passages are fetched; a move to a newer pinned version within the same model family. Worst case: answers get worse across the board, and measurably so.
Route: an ITIL standard change. The CAB approves the class once: what may change, which evaluation set must pass at which score, who signs off, and how rollback works (the previous prompt version from the registry, the previous model ID). After that, individual changes go in without a meeting, and every one is logged against the pre-approval. This is what standard changes were invented for: low risk, well understood, repeatable. The construct is far older than generative AI. Most CABs I meet have simply never been asked to use it for a prompt.
Tier 3: changes to what the system is allowed to do
A new tool or write permission (the system can now create an order, send an email, change a record); a new data source containing personal data; a new user group, especially an external one; a switch to a different model supplier; and any prompt edit that moves or removes a boundary: what the system refuses, what it may promise, when it hands over to a human. Worst case: the system does something it was never approved to do, for someone it was never approved to serve.
Route: a normal change through the CAB, with the privacy officer involved wherever personal data is. And for the minority of systems that fall under the high-risk regime of the EU AI Act, one question a CAB is not used to asking: does this change make us the provider? Article 25 says that a deployer who makes a substantial modification to a high-risk AI system, or changes the intended purpose of a system so that it becomes high-risk, takes on the provider's obligations. The Act has since been amended by the 2026 omnibus, which moved the dates for standalone high-risk systems under Annex III to 2 December 2027, so read the consolidated text before you rely on the detail. The principle, that a large enough change also changes who is responsible, is one you can take into the CAB now.
The one-line prompt edit is the example that convinces CABs. Deleting "never quote a delivery date" is one line, and tier 3. Moving to the next pinned version of the same model is a bigger technical event and, with an evaluation set that passes, tier 2. If your classification puts those the other way round, it is classifying files. Note what this does not mean: the business does not need a ticket for every phrase. Wording inside a pre-approved class is exactly what tier 2 is for. Only the phrases that set a boundary leave it.
What about the changes nobody submits?
The fourth kind of change has no RFC because it has no requester. You cannot approve it. You can only notice it. That makes it the job of monitoring and incident management, not change enablement: a scheduled run of the evaluation set against production, daily or weekly depending on volume, with a threshold that opens a ticket. When the score drops and nothing in your change log explains it, you have found a supplier-side change, and you have the evidence to take to the supplier. Who then picks up that ticket is the question of our earlier essay on who gets paged when an AI system fails.
This is also why the evaluation set keeps turning up in this series. Tier 1 is gated by it, tier 2 is pre-approved against it, and the fourth kind of change is only visible through it. An IT function that does not hold its own evaluation set cannot run any of this, which was the argument of the AI capability you should not outsource.
Does a smaller company need this at all?
No, and I have said so in writing. In a piece on AI strategy for small companies I wrote that a company of thirty should not stand up an AI governance board, and should put one rule on one process instead: who approves an output before it reaches a customer, with a log someone reads at the end of the month. That still holds. This essay is for a different organisation: one where the CAB already exists, meets weekly and has a chair. The question there is not whether to add governance, but which AI changes to keep out of a board that is already there. The line sits roughly where a separate change process already exists. Below it, a log and an evaluation set do the work.
What does ITIL 5 change about this?
PeopleCert launched ITIL (Version 5) in February 2026 and positions it as AI-native. It has published an ITIL AI Governance extension module, and revised practice guides are due in the second half of 2026. Most of what has been written about change in the new version so far concerns AI helping to run change enablement, for example by scoring the risk of a change from historical data. That is useful, and it is a different problem. Using AI to classify your changes does not tell you how to classify changes to AI, and whether the revised guidance will settle that, I cannot tell you yet.
Until it does, the three tiers fit inside ITIL's existing change types without inventing anything. Tier 1 is a data change that stays outside change enablement, tier 2 is a standard change, tier 3 is a normal change. Your CAB does not need a new process. It needs one meeting to agree the classes.
What does this cost the supplier, and why propose it anyway?
In the interest of plain dealing: a classification like this makes my firm's work more visible, not less. Every tier 2 change we make is logged against a pre-approval the client wrote, every tier 3 change waits for the CAB, and the evaluation run shows the client when one of our changes made things worse. A supplier that prefers to ship prompt edits quietly will not propose this.
I propose it because the alternative is worse for both sides. A supplier whose changes are not classified gets blamed for every drift, including the drift it did not cause, and has no log to prove otherwise. In practice this is also most of the distance between an AI demo and AI in production. A demo never needs to know which prompt version produced an answer. A production system does.
Where do you start this month?
- List every AI system in production and, for each one, the four places it can change: prompt, knowledge base, model, supplier.
- Ask your AI supplier where the prompt versions live, and whether every production answer can be traced to a prompt version, an index version and a model ID. If it cannot, that is the first deliverable.
- Take one proposal to the CAB: the three classes, the evaluation threshold for tier 2 and the rollback rule. One agenda item, not a programme.
- Schedule the evaluation run against production, so the fourth kind of change has somewhere to show up.
If you are bringing in a specialist alongside your ICT partner, this classification is part of an AI implementation with us. Either way, it is yours to keep.
We build custom AI and LLM systems that run in production: a clickable MVP by the second call, fixed steps, and you own the code.
AI development agency in the Netherlands →


