The training report says 96 per cent completion and an average satisfaction score of 4.3 out of 5. Both numbers are accurate. Neither tells you whether a single person can do their job.
This is not a new observation, and the framework that names the problem is more than sixty years old and cited in almost every training function — usually while measuring only the parts it says are least informative.
Four levels, two of which get measured
The Kirkpatrick model sets out four levels of training evaluation: reaction, learning, behaviour and results. The order is not accidental — each level is harder to measure and more closely connected to whether the training mattered.
| Level | The question | Usually measured by | A cheap way to do it properly |
|---|---|---|---|
| 1. Reaction | Did they find it useful? | A satisfaction score | One question: “What will you do differently tomorrow?” Blank answers are the finding |
| 2. Learning | Did they acquire the knowledge? | A quiz, if anything | Ask them to complete the task in the system while you watch |
| 3. Behaviour | Are they doing it on the job? | Almost never measured | Unaided completion rate at two and six weeks |
| 4. Results | Did the business outcome change? | Claimed, rarely evidenced | Usually not attributable. Say so rather than inventing it |
Almost all reported training evaluation sits in the first two rows, and level 1 in particular measures something close to enjoyment. A well-delivered, engaging course about the wrong tasks will score highly.
Level 4 deserves an honest word. Attributing a business result to a training intervention, in a programme that also changed the system, the process and the reporting, is rarely defensible. Claiming it damages the credibility of the levels that can be evidenced — and the same attribution problem applies to benefits claimed for adoption generally.
Level 3 is the one worth building
Behaviour is where training either transferred or did not, and it is more measurable than its reputation suggests.
Timothy Baldwin and J. Kevin Ford, reviewing transfer of training in Personnel Psychology, define transfer as requiring both generalisation to the job context and maintenance over a period of time. That second condition is what makes a single post-course measurement insufficient: a test taken on the last afternoon of training measures learning, not transfer.
Three practical measures, in ascending cost:
- Unaided completion rate. Can a named person complete the core task without help, timed? Take it before go-live, at two weeks and at six. The movement is the finding: flat or falling means the work environment is not supporting the behaviour.
- Intervention count. How often did someone need to ask? Already captured if the support model logs it, and it is a direct behavioural measure requiring no extra instrument.
- Error and rework rate on the specific tasks trained. Slower to read but harder to argue with.
The first of these takes half a day per function and produces something a steering committee cannot dismiss, which is more than any completion figure has ever achieved.
Measure the environment, not only the learner
When behaviour does not change, the reflex is to conclude that the training was poor or the people were unwilling. Baldwin and Ford’s model gives a third and usually better explanation: the work environment, which includes supervisory and peer support and the constraints and opportunities to perform the learned behaviour.
So a serious evaluation asks about conditions as well as people:
- Did anyone get time to practise between training and go-live? If not, decay was guaranteed.
- Has the manager discussed the new way with the team since the course?
- Is the old route still available? If the legacy process still works, the new behaviour is optional.
- Does anything measured or rewarded still favour the old way?
Four questions, answerable in a short conversation, and they frequently explain a disappointing level 3 result better than anything about the course does. Markus and Tanis put the general case bluntly, listing no growth of end-user skills after initial training among the characteristic problems of the period after go-live: capability plateaus at whatever the classroom produced unless the environment keeps it moving.
What to report
A defensible training evaluation is short and contains one number most programmes never produce:
- Unaided completion by role, with the date measured. This is the headline.
- Which roles are below the agreed threshold, how many people that is, and what is being done.
- The environmental conditions that did or did not hold, named.
- Completion and satisfaction, reported last, as hygiene rather than as evidence.
Putting completion last is a small act with a large effect. It signals that the training function is being judged on whether people can work, not on whether they attended — which is the standard the rest of the programme is held to.
More on capability evidence in the Change Readiness Hub, or read why training completion and adoption are different measurements.
