How Do I Know Whether AI Adoption Is Actually Working?
You do not know from the dashboard you probably have. Licence counts, training completions and prompt volumes measure what the organisation bought and mandated, not what changed. AI adoption is working when three things are true at once: people return to the tool voluntarily on real work, the output is visibly less generic than before (the Mirror Principle test), and named decisions moved faster or changed because of it. We call the discipline behind this the KPI Inversion: name the goal first, then derive the reading, because a KPI chosen before the goal will flatter the programme and tell you nothing. Stickiness, the voluntary return, is the one adoption signal nobody can produce by mandate. KPMG found only 7 per cent of leaders report having established AI ROI. The other 93 per cent are mostly counting the wrong things.
The numbers you have are the wrong ones
Ask most organisations how their AI adoption is going and you will get a confident answer built from three numbers: seats activated, people trained, prompts sent. Every one of those is a record of activity the organisation itself produced. The licence count went up because procurement bought licences. The training completion rate is high because attendance was required. The prompt volume grew because everyone was told to use the thing. None of it says whether anything about the work has changed.
The honest state of measurement is bleaker than the dashboards suggest. KPMG's global pulse of 2,145 senior leaders found that only 7 per cent report having established AI ROI, while 42 per cent admit they have only partial visibility into what AI is even costing them (KPMG, June 2026). Adoption keeps growing in the same survey. So spend is rising, visibility is partial, and the return is unestablished for more than nine in ten. That is not a measurement gap at the edges of the programme. That is a programme being steered by instruments that read activity and call it value.
I suspect most leaders already sense this. The dashboard says green, and the honest question in the back of the mind stays unanswered: has anyone's actual work changed? The rest of this essay is about how to answer that question with readings that cannot be gamed.
Meet Dave Special: What Changed? Dave is a Group COO whose AI programme is green on every dashboard, until the board asks what business outcome changed. A fictional composite, narrated by Dan Gildoni. Watch the whole series.
Half your people "use AI". Watch how many come back
The single most clarifying thing you can do to an adoption number is put a frequency lens on it. Gallup's quarterly study of 23,717 US employees found that 50 per cent say they use AI in their role. Impressive, until you read the frequency bands underneath: only 13 per cent use it daily, and 28 per cent a few times a week or more (Gallup, April 2026). "Uses AI" and "works with AI" are different populations, and the headline number is built almost entirely from the first one.
The gap between those two figures is where adoption programmes go to be flattered. A seat that logs in once a month counts as adoption on every licence dashboard I have seen, and it is nothing of the kind. What you actually want to know is who returns to the tool of their own accord, on real work, after the launch emails stop. That population is your adoption. Everyone else is your distribution list.
A leader can mandate logins and count prompts. Nobody can mandate the voluntary return.
This is why the daily-use figure is the one worth tracking over time, and why the direction of that figure matters more than its level. A programme where 13 per cent use AI daily and the number is climbing is alive. A programme where 50 per cent "have used" AI and the daily band is flat has already told you its future, whatever the licence dashboard says.
Why the average misleads
Even where value is measured, one number for the whole organisation conceals more than it reveals. Stanford's AI Index reports measured productivity gains of around 14 to 15 per cent in customer support, about 26 per cent in software development, and up to 50 per cent in marketing output, with smaller gains on tasks requiring deeper reasoning (Stanford HAI, 2026). The spread is the finding. An enterprise average across that range describes no team you actually have.
The variation across people is even sharper than the variation across tasks. A randomised experiment by Idan and Anand found that access to generative AI raised average performance, but the gains were highly uneven, and, this is the part worth sitting with, they were not predicted by ability or prior knowledge. They were predicted by what the authors call AI interaction competence: the ability to elicit, filter and verify what the model returns. High-competence participants realised outsized gains. Low-competence participants saw limited or even negative returns (Idan and Anand, 2026). The same tool, in the same role, makes one person faster and another person confidently wrong.
Put those two findings together and the measurement consequence is plain: an adoption average is not a diagnosis. If you must average, average within a team on a named task. Better, stop averaging and start asking where the spread comes from, because the spread is telling you the gap is a reasoning habit, not a licence. That is the same diagnosis we make in why AI training doesn't stick, arrived at from the measurement side.
The KPI Inversion: name the goal, then derive the reading
In coaching sessions with senior leaders, the measurement conversation almost always starts in the same place: "what KPIs should we track for AI?" It is the wrong first question, and we have a name for the correction. The KPI Inversion says: put the word KPI aside and name the goal. KPIs are derivatives. They are what you read once you know what the programme exists to change, and a KPI chosen before the goal will always be one that flatters the programme, because those are the ones the dashboard offers first.
So name the goal in plain words. Not "drive adoption", which is circular, but the thing underneath it: decisions reached faster, preparation that takes a morning instead of a week, work that stops being generic. Then, and only then, ask what would show it. The readings that survive this inversion are almost never the ones on the standard dashboard, and one of them outranks the rest: stickiness. Whether people come back to the tool voluntarily is the honest adoption signal, because it is the only one that cannot be produced by instruction. Every other metric on the board can be moved by a mandate. The voluntary return can only be moved by value.
There is a second discipline hiding inside this one. Reading a return also requires knowing the cost, and KPMG's finding that 42 per cent of leaders have only partial visibility into AI spending (KPMG, June 2026) says the denominator is missing in nearly half of programmes. Judge the programme the way an investor would judge it, as money invested rather than budget allocated. An investor would not accept "seats activated" as a return, and neither should you.
The signals worth reading
Run the inversion on the usual goals of an enterprise AI programme and you end up with a short list of readings. They are harder to collect than a licence count. That is rather the point.
The voluntary return. Of the people who touched the tool in month one, how many are still using it, unprompted, on real work in month three? Read it per team, not as one number, and read the trend before the level. This is the stickiness measure, and it is the closest thing AI adoption has to ground truth. A rising voluntary return means the tool is winning on value. A falling one means the launch is wearing off, whatever the activity metrics say.
The output test. Take real work products from before the programme and now, and ask the question the Mirror Principle asks: is this less generic than it used to be? If the output is generic, the reasoning was generic, and the programme has taught people to produce the same work faster, which is High-Speed Waste wearing a completion certificate. This reading needs a human judge and twenty minutes a quarter. It is worth more than every automated metric combined.
The decision register. Keep a plain list of decisions that moved faster, changed, or got caught before going wrong because of how a team worked with AI. Not estimates, named cases. Five entries in a quarter is a working programme. Zero entries alongside a green dashboard is the clearest possible signal that activity is being mistaken for value. The compression you are looking for is what we call Decision Velocity, and it only ever shows up in specific decisions, never in averages.
I could be wrong about the exact weighting between these, and the right mix shifts by organisation. What I have not seen is a programme that reads all three honestly and stays deluded about whether its adoption is working.
What to stop counting
Retire licence counts as an adoption measure; keep them as a cost line, which is what they are. Retire training completions; they measure attendance, and the evidence that attendance changes behaviour is thin to nonexistent. Retire prompt volume; a thousand vending-machine requests are a thousand pieces of activity, and activity was never the constraint. If a metric can be moved by sending one all-staff email, it is not measuring adoption.
What replaces the retired metrics is a smaller board a leadership team can actually argue about: the voluntary return by team, the output test each quarter, the decision register, and the honest cost line underneath. That board answers the question this page opened with. If the readings are good, the adoption is working and you can say precisely where. If they are bad, you have found out early, from your own data, rather than in the year-three review.
The usual objection is that these readings are harder to collect. They are, and the difficulty is diagnostic: a programme that can only be measured by what it mandated has not yet produced anything voluntary. When the readings do come back bad, the fix is rarely more tools or more training. It is the reasoning habit underneath the keystroke, which is what the Havruta Methodology installs and what whoever owns the mandate, typically the Head of AI Transformation, is actually accountable for. The measurement question and the adoption question turn out to be the same question: has the thinking changed?
Frequently asked questions
How do I know whether AI adoption is actually working?
Not from licence counts or training completions; those measure spend and attendance. AI adoption is working when three things are true: people come back to the tool voluntarily on real work (stickiness), the output is less generic than what the team produced before (the Mirror Principle test), and named decisions moved faster or changed because of it. KPMG found only 7 per cent of leaders report having established AI ROI (KPMG, June 2026), which says most measurement is counting the wrong things.
Why are AI licence counts a bad adoption metric?
Because a licence measures what the organisation bought, not what anyone does. Gallup found that while half of US employees say they use AI in their role, only 13 per cent use it daily (Gallup, April 2026). A seat that is activated and visited once a month is recorded as adoption on a licence dashboard and is nothing of the kind. The honest unit is repeated voluntary use on real work, which is what the KPI Inversion tells you to read instead.
What is the KPI Inversion?
The KPI Inversion is a discipline inside the Havruta Methodology: put the word KPI aside and name the goal first, because KPIs are derivatives of a goal, not a substitute for one. Applied to AI adoption it means asking what the programme exists to change, then reading the one signal that cannot be gamed, whether people return to the tool voluntarily. Stickiness is the honest adoption signal; everything else can be produced by mandate.
What is AI stickiness and why does it matter?
Stickiness is the rate at which people come back to an AI tool of their own accord after the novelty and the mandate wear off. It matters because it is the one adoption measure that cannot be inflated: a leader can mandate logins and count prompts, but nobody can mandate the voluntary return. Gallup's frequency data shows why the lens matters, with 50 per cent of employees reporting some AI use and only 13 per cent using it daily (Gallup, April 2026).
Which KPIs should an AI programme track?
Fewer than most programmes track, and derived from a named goal rather than picked from a dashboard menu. Three readings carry most of the weight: voluntary return rate on real work, output quality against the Mirror Principle (is the work less generic than before), and a register of decisions that moved faster or changed. Cost visibility sits underneath them; KPMG found 42 per cent of leaders have only partial visibility into AI spending (KPMG, June 2026), and a return cannot be read against an unknown cost.
Why do average AI productivity figures mislead?
Because the gains are wildly uneven across tasks and people. Stanford's AI Index reports measured gains from around 14 per cent in customer support to roughly 50 per cent in marketing output (Stanford HAI, 2026), and a randomised experiment found individual gains were predicted not by prior skill but by how well a person elicits, filters and verifies model output, with some users seeing negative returns (Idan and Anand, 2026). An average across that spread describes nobody.
References
- KPMG International. "Growing adoption signals progress as cost visibility and accountability drive AI value." Global AI Quarterly Pulse Survey, Q2 2026, June 2026.
- Gallup. "Rising AI Adoption Spurs Workforce Changes." April 2026.
- Stanford Institute for Human-Centered AI. "The 2026 AI Index Report," Chapter 4: Economy. 2026.
- Idan, L., & Anand, B. "Generative AI and the Productivity Divide: Human-AI Complementarities in Education." arXiv, May 2026.