A copilot writes a reply and you decide whether to send it. An agent sends it.
That is the whole difference, and it is not a difference of degree. Every readiness question changes when the system acts rather than suggests, because the review step — the thing quietly carrying all the risk in current AI deployments — is no longer there.
Most organisations are preparing for agentic AI by evaluating platforms. The harder questions are organisational, they are answerable now, and almost nobody is asking them.
There is already a vocabulary for this, and it is twenty-five years old
In 2000, Raja Parasuraman, Thomas Sheridan and Christopher Wickens published A model for types and levels of human interaction with automation in IEEE Transactions on Systems, Man, and Cybernetics. It decomposes any automated system into four information-processing stages, each of which can be automated to a different degree across ten levels:
- Information acquisition — gathering the raw material
- Information analysis — interpreting it
- Decision and action selection — choosing what to do
- Action implementation — doing it
A copilot automates the first two. It retrieves, summarises, drafts — and hands the third and fourth stages to a human who decides and acts. An agent automates all four.
This framing is more useful than the marketing category, because it stops “agentic” being a binary. The practical question for any deployment is not is this agentic? but which stages have we handed over, and at what level? A system that drafts an order, checks it against credit limits, and queues it for a one-click human approval has automated stages one to three and left four with a person. That is a materially different risk position from one that also places the order, and the two are frequently sold under the same word.
Why the review step was doing more work than anyone credited
Current AI governance rests almost entirely on an assumption that a competent human reads the output before it matters. Checking standards, literacy training, the whole apparatus of “verify before use” — all of it presumes a review point exists.
Remove that and the consequences of the model’s known limitations change character. The BCG field experiment reported by Fabrizio Dell’Acqua and colleagues found that on tasks outside the model’s capability frontier, consultants using AI were 19 per cent less likely to produce correct solutions than those working without it. In a copilot deployment, an outside-the-frontier failure produces a bad draft that costs someone a few minutes. In an agentic deployment, it produces a completed action — an email sent, a record updated, a payment scheduled.
The error rate need not rise at all for the cost of errors to rise by an order of magnitude.
Six questions that are not on the evaluation checklist
1. What may it do without asking? Not what it is capable of — what it is authorised to do unsupervised. This needs to be written as a boundary in business terms: it may issue a refund up to £50; it may not amend a credit limit; it may send an internal message but not an external one. Most agentic deployments inherit the permissions of the account they run under, which is almost never the right answer and is rarely examined.
2. What is reversible, and what is not? Sort the actions into three groups: silently reversible, reversible with effort and embarrassment, and irreversible. An email to a customer is in the second group. A payment is in the third. The boundary of unsupervised authority should sit at the edge of the first group until there is evidence to move it.
3. Who is accountable for an action no person took? Existing accountability structures assume a human decided. When an agent acts within its authorisation and causes harm, the accountable party is whoever set the boundary — and that person needs to know they hold it. Both the NIST AI Risk Management Framework and the EU AI Act put accountability and human oversight at the centre for exactly this reason. Left undefined, accountability settles on the most junior person in the vicinity, which is unjust and also useless as a control.
4. How would you find out it went wrong? This is the question that separates serious deployments from demonstrations. Prevention gets the attention; detection almost never does. If an agent misclassifies a category of case for six weeks, what surfaces that — and how long does it take? A programme that cannot answer this has no feedback loop, and an agent without a feedback loop is an unmonitored process with permissions.
5. Who monitors it, and can a human actually do that? The hardest one, and the best evidenced. Parasuraman and Riley’s work on automation misuse found that over-reliance produces failures of monitoring, and that monitoring degrades precisely when the system is usually right and the workload is high. Asking someone to supervise an agent that is correct 97 per cent of the time, alongside their normal job, is asking them to sustain vigilance against a rare event — a task humans are known to be poor at. “A human is in the loop” is not a control unless that human has the time, information and authority to intervene.
6. What happens to the skill nobody exercises any more? If an agent handles routine cases, people encounter only the exceptions — but competence at exceptions was built by doing the routine work first. Two years in, the population that could confidently judge the agent’s output may no longer exist. This is a slow risk with no obvious detection point, and it is closely related to the expert dependency that already concentrates knowledge in too few people.
| Copilot | Agent | |
|---|---|---|
| Stages automated | Acquisition and analysis | All four, including action |
| Error containment | The human review step | Detection after the fact, if it exists |
| Main capability need | Judging output | Setting boundaries and noticing drift |
| Failure looks like | A bad draft, discarded | A completed action, discovered later |
| Readiness question | Can people judge it? | Can the organisation detect and reverse it? |
What to do before switching anything on
- Write the authorisation boundary in business language, and have the process owner — not the vendor and not IT — sign it.
- Classify actions by reversibility and start unsupervised authority at the reversible end.
- Run it in shadow mode first. Let the agent decide and log what it would have done, without acting, for a defined period. Compare against what humans did. This is the closest thing to a frontier map for actions, and it is cheap.
- Design the detection before the deployment. What sample gets reviewed, by whom, how often, and what threshold triggers a stop. Name the person who can switch it off.
- Resource the oversight honestly. If monitoring is a real job, fund it as one. If it is being added to someone’s existing workload, assume it will not happen when things are busy — which is when it matters.
- Decide how routine competence gets maintained once the routine work is gone. Rotation, deliberate practice, periodic manual handling — something, decided deliberately rather than discovered in year three.
The readiness question underneath
Agentic AI is often framed as a bigger version of the current thing. It is better understood as a transfer of decision rights from people to a system, and organisations have a poor record of doing that explicitly. The decision rights get transferred by configuration, in a project, without anyone experiencing it as a governance decision — which is precisely how policy ends up describing a state the organisation is not in.
The organisations that will handle this well are not the ones with the most capable agents. They are the ones that can answer, for any given process, who authorised what, how a mistake would surface, how quickly it could be undone, and who is accountable when nobody decided anything.
Those questions are answerable today, on the copilot deployment you already have. That is the point of asking them now.
More on trust, oversight and adoption in the AI Adoption Readiness Hub, or use the AI Adoption Readiness Assessment to establish where your organisation stands before decision rights start moving.
