A copilot writes a reply and you decide whether to send it. An agent sends it.

That is the whole difference, and it is not a difference of degree. Every readiness question changes when the system acts rather than suggests, because the review step — the thing quietly carrying all the risk in current AI deployments — is no longer there.

Most organisations are preparing for agentic AI by evaluating platforms. The harder questions are organisational, they are answerable now, and almost nobody is asking them.

There is already a vocabulary for this, and it is twenty-five years old

In 2000, Raja Parasuraman, Thomas Sheridan and Christopher Wickens published A model for types and levels of human interaction with automation in IEEE Transactions on Systems, Man, and Cybernetics. It decomposes any automated system into four information-processing stages, each of which can be automated to a different degree across ten levels:

A copilot automates the first two. It retrieves, summarises, drafts — and hands the third and fourth stages to a human who decides and acts. An agent automates all four.

This framing is more useful than the marketing category, because it stops “agentic” being a binary. The practical question for any deployment is not is this agentic? but which stages have we handed over, and at what level? A system that drafts an order, checks it against credit limits, and queues it for a one-click human approval has automated stages one to three and left four with a person. That is a materially different risk position from one that also places the order, and the two are frequently sold under the same word.

Why the review step was doing more work than anyone credited

Current AI governance rests almost entirely on an assumption that a competent human reads the output before it matters. Checking standards, literacy training, the whole apparatus of “verify before use” — all of it presumes a review point exists.

Remove that and the consequences of the model’s known limitations change character. The BCG field experiment reported by Fabrizio Dell’Acqua and colleagues found that on tasks outside the model’s capability frontier, consultants using AI were 19 per cent less likely to produce correct solutions than those working without it. In a copilot deployment, an outside-the-frontier failure produces a bad draft that costs someone a few minutes. In an agentic deployment, it produces a completed action — an email sent, a record updated, a payment scheduled.

The error rate need not rise at all for the cost of errors to rise by an order of magnitude.

Agentic readiness Evaluating agentic AI for real processes? Book a 20-minute scoping call to work through authorisation, reversibility and who is accountable before anything is switched on.
Book a 20-minute scoping call

Six questions that are not on the evaluation checklist

1. What may it do without asking? Not what it is capable of — what it is authorised to do unsupervised. This needs to be written as a boundary in business terms: it may issue a refund up to £50; it may not amend a credit limit; it may send an internal message but not an external one. Most agentic deployments inherit the permissions of the account they run under, which is almost never the right answer and is rarely examined.

2. What is reversible, and what is not? Sort the actions into three groups: silently reversible, reversible with effort and embarrassment, and irreversible. An email to a customer is in the second group. A payment is in the third. The boundary of unsupervised authority should sit at the edge of the first group until there is evidence to move it.

3. Who is accountable for an action no person took? Existing accountability structures assume a human decided. When an agent acts within its authorisation and causes harm, the accountable party is whoever set the boundary — and that person needs to know they hold it. Both the NIST AI Risk Management Framework and the EU AI Act put accountability and human oversight at the centre for exactly this reason. Left undefined, accountability settles on the most junior person in the vicinity, which is unjust and also useless as a control.

4. How would you find out it went wrong? This is the question that separates serious deployments from demonstrations. Prevention gets the attention; detection almost never does. If an agent misclassifies a category of case for six weeks, what surfaces that — and how long does it take? A programme that cannot answer this has no feedback loop, and an agent without a feedback loop is an unmonitored process with permissions.

5. Who monitors it, and can a human actually do that? The hardest one, and the best evidenced. Parasuraman and Riley’s work on automation misuse found that over-reliance produces failures of monitoring, and that monitoring degrades precisely when the system is usually right and the workload is high. Asking someone to supervise an agent that is correct 97 per cent of the time, alongside their normal job, is asking them to sustain vigilance against a rare event — a task humans are known to be poor at. “A human is in the loop” is not a control unless that human has the time, information and authority to intervene.

6. What happens to the skill nobody exercises any more? If an agent handles routine cases, people encounter only the exceptions — but competence at exceptions was built by doing the routine work first. Two years in, the population that could confidently judge the agent’s output may no longer exist. This is a slow risk with no obvious detection point, and it is closely related to the expert dependency that already concentrates knowledge in too few people.

 CopilotAgent
Stages automatedAcquisition and analysisAll four, including action
Error containmentThe human review stepDetection after the fact, if it exists
Main capability needJudging outputSetting boundaries and noticing drift
Failure looks likeA bad draft, discardedA completed action, discovered later
Readiness questionCan people judge it?Can the organisation detect and reverse it?

What to do before switching anything on

The readiness question underneath

Agentic AI is often framed as a bigger version of the current thing. It is better understood as a transfer of decision rights from people to a system, and organisations have a poor record of doing that explicitly. The decision rights get transferred by configuration, in a project, without anyone experiencing it as a governance decision — which is precisely how policy ends up describing a state the organisation is not in.

The organisations that will handle this well are not the ones with the most capable agents. They are the ones that can answer, for any given process, who authorised what, how a mistake would surface, how quickly it could be undone, and who is accountable when nobody decided anything.

Those questions are answerable today, on the copilot deployment you already have. That is the point of asking them now.

More on trust, oversight and adoption in the AI Adoption Readiness Hub, or use the AI Adoption Readiness Assessment to establish where your organisation stands before decision rights start moving.

Ritvars Mētra

Ritvars Mētra

Founder of ReadinessCompass

Ritvars Mētra is the founder of ReadinessCompass, where he develops practical tools for understanding and managing organisational change complexity. His work focuses on adoption readiness, stakeholder analysis, and evidence-based change management for large-scale software and AI implementations.

View full profile →