The prompt engineering workshop went well. Ninety minutes, good attendance, people left with a one-page cheat sheet of patterns — give it a role, give it context, ask for a specific format, iterate.
Six weeks later, usage has not moved in any function that matters, two people have pasted customer data into a consumer tool, and a manager has forwarded a confidently written summary containing a figure that does not exist anywhere in the source document.
None of that is a prompting failure. Every one of those people could write a competent prompt. What they could not do was judge the output, recognise where the boundary was, or decide when the tool should not have been used at all — and none of that was on the cheat sheet.
There is now an obligation, and it does not mention prompting
For any organisation operating in the EU this stopped being a purely developmental question on 2 February 2025, when Article 4 of the EU AI Act came into application. It requires providers and deployers of AI systems to take measures to ensure a sufficient level of AI literacy among their staff and other people operating or using AI systems on their behalf.
The definition is the interesting part. AI literacy is framed as the skills, knowledge and understanding that allow people to make an informed deployment of AI systems, and to gain awareness of the opportunities and risks of AI and the possible harm it can cause. Measures must take into account people’s technical knowledge, experience, education and training, and the context in which the systems will be used.
Three things follow from that wording. It covers deployers, not just builders — if you bought a tool and rolled it out, it applies to you. It is explicitly about risk and harm awareness, not productivity technique. And it is contextual, which means a single generic course delivered identically to everyone is a poor fit for the obligation by design.
Notably, the obligation does not require guaranteeing any particular level of literacy in any individual. It is about the measures you take. That is a lower bar than it first appears, and also an awkward one: “we ran a webinar” is a measure, and so is a serious role-based programme, and only one of them will hold up if something goes wrong.
What the research says literacy consists of
The most cited academic treatment is Duri Long and Brian Magerko’s What is AI Literacy? Competencies and Design Considerations, presented at CHI 2020. They synthesise interdisciplinary literature into seventeen core competencies, and define AI literacy as the set of competencies that enables people to critically evaluate AI, communicate and collaborate effectively with AI systems, and use AI appropriately across different areas of life and work.
Critically evaluate. Collaborate effectively. Use appropriately. Prompt craft sits inside the second of those and touches neither of the others — and the failures that hurt organisations come almost entirely from the first and third.
| Prompt training | AI literacy | |
|---|---|---|
| Teaches | How to get a better output | Whether to accept the output, and whether to have asked at all |
| Assumes | The task is appropriate for the tool | Nothing; appropriateness is the first judgement |
| Failure it prevents | A weak answer | A confident wrong answer acted on, or a disclosure that should not have happened |
| Transfers across tools | Partly — patterns shift with each model | Yes — the judgement is tool-independent |
| Can be tested by | Whether someone produces a usable prompt | Whether someone catches a plausible error they were not warned about |
The competency that matters most is calibrated trust
If you teach one thing, teach this, and there is thirty years of human factors research behind it.
In Human Factors in 1997, Raja Parasuraman and Victor Riley published Humans and Automation: Use, Misuse, Disuse, Abuse, which remains the standard vocabulary for how people get their relationship with automation wrong. Two of their categories describe almost every AI failure now occurring in offices:
- Misuse — over-reliance on automation, producing failures of monitoring and decision biases. The manager who forwards the summary without checking the figure. Monitoring degrades further when workload is high and when the system is usually right, which is precisely the condition modern language models create.
- Disuse — neglect or underuse of automation, classically caused by false alarms. The analyst who tried it twice, caught it inventing a citation, and now refuses to use it for anything — including the tasks where it would be reliable and valuable.
Most organisations have both populations simultaneously and treat neither, because the training on offer addresses a third thing entirely. Misuse and disuse are not opposite problems requiring opposite messages; they are the same missing competency, which is the ability to judge this output on this task rather than holding a fixed global attitude towards the tool.
Parasuraman and Riley also name a failure that belongs to designers and managers rather than users — automation deployed without regard for the human role around it. Rolling out a copilot into a process nobody has redesigned, and then attributing the disappointing result to user resistance, is that failure exactly. It is also why working with probabilistic output is a genuinely new readiness demand rather than a familiar training exercise.
Four things employees actually need
1. How this tool fails, specifically. Not “AI can be wrong”, which everyone already knows and nobody acts on. People need the characteristic failure modes of the system in front of them: it invents plausible references; it is confident about arithmetic it has not done; it will summarise a document it only partially read; it reflects the framing of the question back at you. Failure modes are learnable and they are what makes checking efficient rather than paranoid.
2. Where the boundary is, in their own words. Not a policy document. A short, role-specific list of what must never be pasted in, expressed in the vocabulary of their actual work — “a candidate’s name and their interview notes”, not “personal data as defined under applicable regulation”. The two people who pasted customer data had almost certainly read the policy.
3. When not to use it at all. The judgement nobody teaches. Tasks where the cost of a subtle error exceeds the time saved; tasks where the reasoning matters more than the answer; tasks where a person is entitled to a human decision. Making this explicit protects people, and it is also the honest counterweight to enthusiasm.
4. A checking method proportionate to the stakes. “Verify the output” is not actionable. What is actionable: for internal drafts, read it once for sense; for anything containing a number, check every number against source; for anything leaving the organisation, a named second reader. Three tiers, written down, and the rule stated as a property of the task rather than a property of the person’s diligence.
Why generic courses underperform
A finance analyst and a recruiter face different risks, different boundaries and different failure modes. The analyst’s danger is a number that looks right. The recruiter’s is a judgement that quietly encodes a bias and touches a person’s livelihood. A single course covering both at the level of generality that permits it will be true, forgettable and behaviourally inert.
This is also what the Article 4 wording pushes towards, since it requires measures to account for the context in which the systems are used. Practically: one short common module on how these systems fail, then role-specific sessions built around three or four real tasks that group actually does, using their own material.
It is the same principle that separates training from capability everywhere else — being shown something is not the same as being able to do it under pressure.
How to know whether it worked
Attendance and completion tell you nothing here, for the reasons that make them weak criteria in any readiness context: they measure whether people were processed.
The test that works is a planted error. Give people a realistic AI-generated output for their own role, containing two defects — one obvious, one plausible and consequential — and ask them to use it as they normally would. What proportion caught the plausible one? That single number is a better literacy measure than any survey, it is directly comparable across functions and over time, and it can be run in twenty minutes.
Two other signals worth watching:
- Whether people can say where the boundary is without looking it up. If they cannot state it, the policy exists but the literacy does not.
- Whether disuse is falling. The population that abandoned the tool after one bad experience is the clearest evidence that failure modes were never taught — and they are usually invisible in adoption dashboards, which count active users rather than absent ones.
Managers need a different curriculum
One group is routinely trained as though they were users when their problem is different. Managers do not mainly need to operate the tool. They need to be able to answer the questions their teams will actually ask: is it acceptable to use this here, does using it count against me, what happens if it is wrong and I did not catch it, and are my targets going to move because of this.
A manager who cannot answer those confidently will default to discouraging use in public and tolerating it in private, which produces precisely the concealed, unmeasurable adoption pattern that shadow AI represents.
The distinction worth holding
Prompt training makes people faster at getting output. AI literacy makes them safe to give output to. Only one of those is a readiness condition, only one is now a legal obligation for deployers in the EU, and only one survives the next model release.
The organisations that will get value from this are not the ones whose staff can write elaborate prompts. They are the ones where a mid-level employee can look at a fluent, well-formatted, entirely wrong answer and say so — and where saying so is treated as competence rather than obstruction.
More on capability and trust in the AI Adoption Readiness Hub, or use the AI Adoption Readiness Assessment to see where your organisation stands before committing to a training approach.
