Insights

We Rolled Out Copilot and Nothing Changed. What Do We Do Now?

In short

If Copilot went out to everyone and nothing about the work changed, your rollout was not a failure. It was the normal result. MIT's NANDA research found tools like ChatGPT and Copilot widely adopted across enterprises while lifting individual convenience rather than profit-and-loss performance, and Microsoft's own 2026 research shows the value concentrating in a small minority who work with the tool differently. The problem is not the licence, the model, or your people's appetite. It is that deployment changed what people have access to and left untouched how they reason with it, the pattern we call the Vending Machine reflex. Two moves change the result: anchor the tool to one real bottleneck per team, and install a reasoning habit on live work. Then read the one signal that cannot be gamed, whether people come back voluntarily.

On this page
  1. Your rollout was normal, and that is the bad news
  2. The vendor's own numbers
  3. What actually happens after rollout
  4. The diagnosis: the machine is used as a vending machine
  5. The two moves that change the result
  6. The honest test before you touch the licences
  7. Frequently asked questions
  8. References
Pencil-sketch illustration: an executive at a desk between two identical stacks of documents, his AI assistant asking What are you trying to decide?
Same desk, same stacks. The question on the screen is where the change starts.
01 · The company you keep

Your rollout was normal, and that is the bad news

The scene is the same in most large organisations I speak to. The licences went out with some fanfare. There was a launch, a training wave, a channel for tips. Eighteen months later the honest people in the room admit that the work looks the way it looked before, and somebody senior is asking what, exactly, the spend bought.

Here is the uncomfortable comfort: this is the majority outcome, measured. MIT's NANDA research put it plainly: "Tools like ChatGPT and Copilot are widely adopted. Over 80 percent of organizations have explored or piloted them, and nearly 40 percent report deployment. But these tools primarily enhance individual productivity, not P&L performance" (MIT NANDA, 2025). The same research found just 5 per cent of integrated AI pilots extracting real value, and named the cause: not model quality, but tools that do not learn the organisation's context and workflows that did not change to meet them. Connecting the tool to an AI knowledge base fixes access to that context; it does not, by itself, change the decision.

So the first thing to do is stop treating your rollout as a local failure. It was a standard deployment producing the standard result. The useful question is not "what went wrong here" but "what do the few who get a different result do differently", and on that question the evidence is unusually clear, because the vendor publishes it.

02 · The vendor's numbers

Microsoft's own research says the gap is you, not the tool

Microsoft's 2026 Work Trend Index, built from trillions of anonymised Microsoft 365 signals and a survey of 20,000 AI-using workers, is worth reading closely, because it is the vendor describing its own adoption problem. Three findings stand out.

First, the value concentrates hard. The advanced users Microsoft calls Frontier Professionals, the ones running multi-step work with the tools rather than one-off requests, are just 16 per cent of the AI users surveyed (Microsoft, 2026). Everyone else is using the same licence at a fraction of its depth. Second, the cause is organisational: Microsoft found that factors like culture, manager support and talent practices account for more than twice the reported AI impact of individual factors like mindset (67 per cent against 32). And third, only 26 per cent of AI users say their leadership is clearly and consistently aligned on AI.

Hold those three together and the shape of the problem changes. The tool works. A minority works with it in a way that compounds. And what separates that minority is not talent or licence tier but the surrounding discipline, which is something leadership either builds or does not. When the vendor's own data says the binding constraint is on your side of the table, I would take the vendor at its word.

03 · The arc

What actually happens after a rollout

The post-rollout arc has now been measured longitudinally, and it is not a plateau. It is a narrowing. A study that followed a Microsoft 365 Copilot pilot at a US state transport department found perceived usefulness fell over the pilot, from 3.85 to 3.62 on a five-point scale, and usage contracted rather than deepened: data and chart generation dropped by nearly half, presentations likewise, while summarising and drafting held steady (Shoghli et al., 2026). People tried the harder uses, got burned on accuracy-sensitive work, and retreated to the shallow end. Fraunhofer's study of a 550-licence rollout found the same shape from the other side: the tool rated most useful for structured, text-based tasks, with the researchers pointing at routinisation and the need for role-specific embedding rather than any property of the model (Schmidt et al., 2026).

A UK government evaluation of a 1,000-licence Copilot pilot, published back in August 2025, had already reached the blunt version of this conclusion: it did not find evidence that the time savings led to improved productivity, and colleagues outside the pilot noticed no difference in those inside it (UK Department for Business and Trade, 2025). A full year on, the pattern it described is still the default arc.

Notice what none of these studies found: resistance. People used the tool. Many liked it. The work still did not change, because the way people engaged the machine did not change. Which brings us to the diagnosis.

04 · The diagnosis

The machine is being used as a vending machine

Watch how most people actually use their assistant and you will see the same transaction on a loop: summarise this document, draft this email, tidy this deck. Insert request, take output, repeat. The machine is never given a role, never told the real goal, never allowed to question whether the request is even the right one. We call this Vending Machine vs Thinking Partner, and it is the central diagnostic for everything this page describes. A vending machine dispenses. It does not ask why you are hungry.

The rollout made this worse in a way nobody intended. Because the barrier to entry is zero, everyone could start using the tool immediately, with no training in how to reason with it, and so everyone defaulted to the transaction. That is the Zero-Barrier Trap: the barrier to entry is zero, and the barrier to value is business training. The organisation cleared a barrier that did not exist and left standing the one that did.

The licence changed what people have access to. It could not change what they ask.

This is why the usage narrowed in the studies above. Vending-machine use finds its level at the tasks a vending machine is good at, which is exactly the summarise-and-draft shallow end the longitudinal data shows. Deeper value needs a different opening move, and the opening move is teachable.

05 · The fix

The two moves that change the result

Everything I have watched work in real organisations reduces to two moves, made in this order.

Move one: anchor the tool to a named bottleneck. Stop deploying to everyone in general and start anchoring to one problem per team, the specific piece of work that costs that team the most pain each week. The unit of AI value is the human bottleneck, not the use case, and general availability is precisely how a tool ends up owned by nobody and pointed at nothing. A team that solves its own named problem with the machine does not need a champions programme afterwards. The result recruits the next team.

Move two: install the reasoning habit on live work. On that named problem, replace the transaction with the 4-Lines: give the machine a specific expert persona, state the real goal, make it question you before it answers (the Flip), and run the exchange step by step. This is done on the team's actual work, in the room, not in a training environment, because the evidence on why AI training fails to change behaviour is unambiguous about the difference. The habit is small enough to learn in an afternoon and big enough that Microsoft's 16 per cent are, in effect, the people who already have some version of it.

Neither move requires a new tool or another licence. Both work on whatever the organisation already deployed, which is rather the point: the spend is sunk, and these are the moves that make it pay.

06 · Before you cut

The honest test before you touch the licences

At some point in this conversation somebody proposes cancelling the licences, and I understand the instinct. Resist it for one quarter. Cancelling first teaches you nothing except that unused licences cost money, and the same failure will repeat under the next tool, because the habit travels with the people.

Run the honest test instead. Pick two teams. Anchor each to one named bottleneck, install the habit on live work, and give them eight weeks. Then read the signals that cannot be gamed, the ones we set out in how to know whether AI adoption is actually working: the voluntary return, the output against the Mirror Principle, and the decisions that moved. Compare against a team you left alone. If nothing moves under those conditions, you have learned something real about your organisation. In my experience that is not what happens. What happens is the two teams stop being an experiment and start being the argument.

The rollout put the machinery on every desk. What it could not install is the discipline of thinking with it, and that is not a criticism of the rollout. It is the definition of the remaining work, and the remaining work is what the Havruta Methodology exists to do. Whoever owns this in your organisation, the CIO or the Head of AI Transformation, the sequence is the same: one bottleneck, one habit, one honest reading.

07 · Frequently asked

Frequently asked questions

We rolled out Copilot and nothing changed, what do we do now?

Stop treating it as a tool problem and start treating it as a reasoning problem. The evidence says the rollout was normal: MIT NANDA found tools like ChatGPT and Copilot are widely adopted yet primarily lift individual convenience, not profit-and-loss performance. Two moves change the result: anchor the tool to one real bottleneck per team rather than general availability, and install a reasoning habit (the 4-Lines) on live work so people stop using the machine as a vending machine. Then measure the voluntary return, not the licence count.

Why does Copilot not improve productivity?

Mostly because it is pointed at the same work in the same way. A longitudinal study of a government Copilot pilot found perceived usefulness fell over the pilot and usage narrowed to simple text tasks such as summarising and drafting (Shoghli et al., 2026), and Fraunhofer found the tool valued mainly for structured, text-based activities. The machine does what it is asked. When what it is asked is the old work, the output is the old work, faster. That habit, not the model, is the constraint.

How do executives use Copilot effectively?

By changing the opening move, not the tool. An executive gets value from Copilot the same way they get value from any AI: give it a specific expert persona, state the real goal rather than the surface task, make it question you before it answers (the Flip), and work step by step. That is the 4-Lines, and it turns the assistant from a faster typist into a thinking partner. Microsoft's own research finds the advanced minority who work this way, about 16 per cent of AI users, capture disproportionate value (Microsoft, 2026).

AI licences bought, adoption flat, what is the fix?

The fix is business training, not more technical enablement, because the barrier to entry is zero and the barrier to value is knowing how to reason with the machine on real work. We call this the Zero-Barrier Trap. Anchor each team to one named problem it owns, install the 4-Lines habit on that problem, and let the result recruit the next team. Microsoft's 2026 Work Trend Index found organisational factors account for more than twice the reported AI impact of individual factors, so the fix is a leadership move, not a user move.

Should we cancel our Copilot licences?

Probably not, and the licence question is the wrong question. The spend is rarely the problem; the missing return is, and it stays missing under any tool if the way people engage it stays transactional. Before cutting, run one honest test: pick two teams, anchor each to a real bottleneck, install the reasoning habit, and read the voluntary return after eight weeks against a control team. If nothing moves under those conditions, you have learned something real. Cancelling first teaches you nothing except that unused licences cost money.

What is Vending Machine vs Thinking Partner?

Vending Machine vs Thinking Partner is the central diagnostic of the Havruta Methodology. Vending-machine use is transactional: insert a request, take the output, repeat, with the machine never questioning the request. Thinking-partner use is paired reasoning: the machine gets a persona and a real goal, challenges the request before answering, and works step by step. The same licence supports both. Which one an organisation gets is decided by habit, which is why deployment alone changes nothing.

The licences are already paid for. The habit that makes them pay is what we install.