Home / Insights / AI Agent Autonomy: Treat It Like a New Hire
Insights

AI Agent Autonomy: Treat It Like a New Hire

Summarize with AI Prompt copied — paste it into the chat

Most Dutch businesses giving an AI agent real work this year skip a step no one would skip with a new hire: a trial period with limited authority, checked output, and a clear moment where trust is extended or withdrawn. Treat a new AI agent the way Dutch employment law treats a new employee — a proeftijd, not a blank cheque — and most of the governance failures I see in first conversations simply stop happening.

I've had a version of the same conversation a dozen times this year. A business owner wants an agent that reads incoming email, checks stock in Exact Online or AFAS, and drafts — or sends — a reply. The technology to do this well already exists. The part that goes wrong almost never breaks in the model. It breaks in the handover: someone grants full mailbox and ERP access on day one, with no plan for what happens when the agent gets something wrong, because nobody wrote down what "wrong" would look like in advance.

The size of business where this comes up most isn't the sceptics and it isn't the early adopters — it's the 20-to-50-person company that just watched a competitor cut its response time in half and wants the same result by next quarter. That urgency is legitimate. It's also exactly the condition under which people skip the proeftijd instinct entirely, because slowing down to define scope feels like the opposite of the speed they were promised.

Why AI agent projects fail: too much trust, or too little

In most first conversations about agents, I hear one of two instincts. The cautious owner wants to lock everything down — read-only access, a human checking every output, no exceptions — and then wonders six months later why the agent never saved anyone any time. The confident owner wants to skip straight to full autonomy, because the demo looked flawless, and then spends a Friday afternoon explaining to a customer why the agent sent an apology email to the wrong account.

Gartner's research team named this pattern precisely in May 2026: organisations that apply the same governance to every agent, regardless of what it's actually allowed to touch, run into one of two failure modes — over-restriction that kills adoption, or under-restriction that turns a small mistake into an incident. Gartner's fix is a graduated model: classify agents by autonomy level, and match the amount of oversight to what the agent can actually do, not to how nervous the room is about AI in general. That's precisely the model that already exists, culturally, in how the Netherlands hires people.

The proeftijd as a way of thinking about agent autonomy

Dutch employment law gives every new hire a proeftijd — a probationary period, capped at one month for a fixed-term contract under two years, two months for an indefinite contract or a longer fixed term, and not allowed at all for a contract of six months or less. During it, either side can walk away without notice and without giving a reason. It exists because a CV and an interview can't fully predict how someone performs with real customers, real deadlines, and real access to the systems that run a business. Nobody hands a new employee the company bank login on their first Monday, and nobody should hand an AI agent full write access to a mailbox or an ERP on its first day either.

I use that same instinct — deliberately, not as a metaphor I reach for once and drop — when a client wants to put an agent to work inside their business. This is a Dutch legal concept, not a Belgian one; Belgium's eenheidsstatuut reform abolished the standard proefbeding for ordinary employment contracts back in 2014, replacing it with graduated notice periods instead. If you're reading this in Flanders and the idea still resonates, the proeftijd habit works fine as a mental model — you just won't find it in your own arbeidsrecht.

The parallel isn't decorative. It forces four honest questions before an agent goes live: what is this agent actually allowed to do, who checks its work and how often, what happens the moment it gets something wrong, and who has the authority to end the trial. Most agent rollouts I see skip straight past all four and go directly to "does it work in the demo."

None of this is unique to agents, either. I ask nearly the same four questions before any new piece of automation touches a live system — the difference with an agent is that the model itself decides, case by case, what "doing the job" looks like, so the four answers can't live only in a kickoff document nobody rereads afterwards. They have to live in what the agent is actually, technically, allowed to touch.

The four levels of autonomy I actually use

Gartner's framework sorts agents into four levels, and I map nearly every agent conversation I have onto it before we write a line of code or a system prompt.

  • Level 1 — Observe: read-only access to a defined data source, output visible only to the person who asked. This is where every new agent starts in my proeftijd, no exceptions, no matter how good the pilot looked.
  • Level 2 — Advise: the agent drafts — an email, a report, a work order — but a person reads it before anything leaves the building. This is where most Dutch small-business agents should live for months, not weeks.
Pull quote: Give your AI agent a proeftijd, not a blank cheque. — Crux Digits
  • Level 3 — Act with approval: the agent can send the email or update the record, but only after an explicit yes from a named person, every single time, with a log of who approved what.
  • Level 4 — Act autonomously: the agent acts inside guardrails without per-action approval, and a person reviews exceptions and aggregate outcomes instead. Very few of the SME agents I've seen earn this level within their first year, and that's a feature, not a failure.

Picture a new receptionist in their first week. You wouldn't hand them the master keys and the alarm code before they'd answered a single call under supervision — you'd have them observe calls, then draft messages for someone else to send, then handle simple requests with a colleague nearby, and only later work the front desk alone. An agent that goes straight from a sales demo to Level 4 has skipped every one of those weeks, and it shows in exactly the same way an under-trained new hire shows: not in raw ability, but in the judgement calls nobody tested for.

The promotion from one level to the next should feel exactly like a proeftijd ending well: a specific date, a specific person deciding, and a specific, small number of failures over a specific number of weeks used as the bar — not a feeling that "it's probably fine by now."

What I check before I hand an agent the keys

Before any agent of mine moves past Level 2, I want four things written down, the same way I'd want a job description written down before someone starts:

  • A scope as narrow as a job description: not "handles customer service" but "drafts replies to stock-availability questions from the webshop address, nothing else."
  • A visible log of every action the agent took, in a place a non-technical person can actually read — not a database table only a developer opens.
  • A rollback plan: what happens in the ten minutes after the agent sends the wrong thing, and who is responsible for those ten minutes.
  • A named owner — not "the AI system", a person whose job includes deciding when the proeftijd ends, same as a manager signs off on a new hire.

I also want to see the agent handle the request that isn't in the happy path — a stock question about a product that's been discontinued, an email written in a dialect the model wasn't tuned on — before a single action moves from draft to sent. A proeftijd that only tests the easy cases isn't a proeftijd; it's a formality.

This is also where a written AI usage policy earns its cost — and where an ongoing AI-beheer retainer, someone checking that log every week rather than just at launch, tends to pay for itself many times over. Most of the businesses I meet that skip this step aren't being reckless — they simply haven't been asked the question in a form that made them stop and answer it. Writing it down is the whole exercise.

When an agent doesn't survive its proeftijd

Gartner predicts that by 2027, 40% of enterprises will demote or decommission an autonomous AI agent because a governance gap only became visible after something went wrong in production. I expect the real number for the Dutch mkb to run higher, not lower — smaller teams have less slack to absorb the fallout of an over-trusted agent, and less appetite to publicly admit an experiment didn't work.

I've sat across the table from owners halfway through that reversal, and the conversation is never really about the technology. It's about admitting, out loud, that the scope was wrong on day one — which turns out to be a much smaller thing to admit than most people expect, once the proeftijd framing makes it normal rather than embarrassing.

That's not a reason to avoid agents. It's a reason to build the exit into the plan from day one, the way a proeftijd already builds in an exit for a new hire: no drama, no lengthy process, just a decision that the fit wasn't right and a return to the previous, manual way of doing that one task while you try again with a narrower scope. The agents I've seen quietly get switched back to Level 1 or 2 after a rough month tend to come back successfully later. The ones that get switched off entirely, without anyone reviewing why, usually get tried again from scratch a year later — at Level 4, on day one, by someone who's forgotten the lesson.

What this means for your AI policy — and a deadline you just passed

As of 2 August 2026, Article 50 of the EU AI Act's transparency obligations apply: if a customer or employee is interacting with an AI system, they need to be told. That requirement sits comfortably inside the proeftijd model — a Level 2 or Level 3 agent that a person reviews before anything reaches a customer is, by construction, easy to be transparent about, because a human is already in the loop. An agent that was rushed straight to Level 4 without that habit is the one that hands a business a compliance problem on top of an operational one. Our AI Act check walks through what the August deadline actually requires for a business your size, separate from the high-risk obligations that were pushed back to December 2027.

The bigger number behind all of this: CBS's most recent AI monitor puts Dutch AI use at 22.7% of businesses with 10 or more employees in 2024, nearly double the year before, but still only 17.8% among businesses with 10 to 19 staff — the size band where most of my agent conversations actually happen. Adoption is accelerating faster than governance habits are, which is exactly the gap a proeftijd is built to close. The full picture on Dutch adoption is here. An agent earns full trust the same way a new colleague does: not on day one, and not because the interview — or the demo — went well, but because the work held up once someone was actually checking it.

Frequently asked questions

How long should an AI agent's probation period last?

Long enough to see it handle real edge cases, not just the happy path — for most Dutch SME agents that's four to eight weeks at Level 2 before any promotion, roughly mirroring the one-to-two-month proeftijd most new hires get under Dutch law.

Who is liable if an AI agent makes a mistake?

The deploying business, in essentially every case I've encountered — not the model vendor. That's exactly why a named human owner and an approval log at Level 3 matter more than the contract you signed with the AI supplier.

Do I have to tell customers they're talking to an AI agent?

Yes. Since 2 August 2026, Article 50 of the EU AI Act requires it — the disclosure duty applies regardless of company size or how autonomous the agent is.

How many AI agents does a small Dutch business actually need?

Usually one, doing one task well, before a second is worth the governance overhead. The businesses that try to run five agents at once are almost always the ones that can't tell me which single agent is actually saving time.

Can an AI agent be 'let go' if it doesn't work out?

Yes, and it should be treated that plainly — switch it back to a lower autonomy level, or off entirely, without treating the reversal as a verdict on your whole AI strategy.
Our AI services Hire an AI consultant AI automation AI agents AI implementation Pricing

Want any of this applied to your business?

We turn these concepts into working tools — grounded, safe and measurable. Start with a free consultation.

Book a free consultation →