A readiness survey asks people what they believe about their own capability. A transaction log records what they did.
Both are useful and they answer different questions. Most programmes use only the first, and then wonder why the readiness score and the operational reality diverged.
Why self-report has a ceiling
Self-report is not unreliable because people lie. It is limited because of how the instrument works. Norbert Schwarz’s Self-reports: How the questions shape the answers establishes that reports of behaviour and attitudes are strongly and systematically influenced by question wording, format and context.
There is also the disclosure problem. Microsoft and LinkedIn’s 2024 Work Trend Index found 52 per cent of AI users reluctant to admit using it for their most important tasks. Where an honest answer carries a cost, self-report understates — and it understates most in the situations you most need to understand.
Logs have neither problem. They also have their own, which is that they record what happened and not why.
Six things worth extracting
| Signal | What it indicates |
|---|---|
| Time to complete, as a distribution | Not the average — the spread. A long tail means a subgroup is struggling that the mean conceals |
| Reversals, cancellations and corrections | Rework. The closest thing to a quality measure available without extra instrumentation |
| Path deviation | People reaching an outcome by an unintended route. Almost always a design problem, and the origin of workarounds |
| Feature abandonment | Something used in week one and not in week six. Usually means it did not work as expected and nobody reported it |
| Concentration | A small number of user IDs performing a large share of a transaction type — the signature of expert dependency |
| Timing patterns | Work batched immediately before a review, or done outside normal hours, indicating the system is not part of the daily flow |
Concentration is the most under-used of these and among the most valuable. It answers a question no survey will: whether capability actually spread, or whether three people are quietly carrying a function.
What logs cannot tell you
Why. A rising correction rate could be capability, data quality, a design flaw or a genuine increase in complex cases. The log identifies where to look; it does not diagnose.
Whether the outcome was right. A transaction completing cleanly is not evidence it was the correct transaction. Speed without a quality measure is the most dangerous number in an adoption dashboard.
What is happening outside the system. The spreadsheet does not appear in the log. Neither does the phone call that resolved the exception, nor the colleague who completed the transaction on someone else’s behalf — which will read as that colleague’s competence.
So the workable method is to use logs to locate and surveys or observation to explain. A log that shows a long completion tail in one team tells you where to spend the half day watching people work.
The constraint most programmes discover late
Analysing system logs at individual level is employee monitoring, and in much of Europe that is legally constrained rather than merely sensitive.
In Germany the works council has a co-determination right over the introduction of technical equipment suitable for monitoring conduct or performance — and case law holds that mere suitability is enough, regardless of intent. That standard captures almost every system a programme deploys, and it applies to how you analyse the data as well as to deploying the system.
Three practices keep this straightforward. Analyse at team or role level rather than individual, which answers the adoption question anyway. Apply the same reporting thresholds you would to any small-group readiness data. And agree the analysis with employee representatives in advance rather than after someone notices — the questions they ask are the ones you should have answered before starting.
Use both, and expect them to disagree
The most informative moment in adoption measurement is when the survey and the log conflict. High reported confidence with a long completion tail means people do not know what they do not know. Low reported confidence with clean logs usually means an anxious team that is performing fine and needs telling so.
Neither instrument is the truth. The disagreement between them is where the finding is.
More on adoption measurement in the Change Readiness Hub, or read what to measure after go-live.
