Back to blog
Last updated July 2026·Rozbeh Karimi

How to Measure the ROI of an AI Workshop

The first month after a workshop is when you can tell whether it worked, and almost nobody measures it — so almost nobody finds out. They measure whether people enjoyed the day, which is a different question with a reliably flattering answer.

Three numbers tell you what actually happened: how many hours people say they saved, what changed in their output, and what share of the team used AI at all this week. The third one predicts the other two. This is how to collect all three without building a reporting apparatus nobody maintains.

Why satisfaction scores mislead you

Nearly every workshop ends with a feedback form, and nearly every feedback form comes back positive. People generally enjoy a well-run day away from their inbox, and they will tell you so. That number is close to useless for deciding whether to do it again.

Enjoyment and adoption are barely correlated. I have seen sessions with glowing feedback where nothing changed, and sessions people found hard going that shifted how a whole department worked. The four ways workshops fail are all visible in the numbers below. The only honest test is behavioural: what are people doing four weeks later that they were not doing before?

The three metrics that matter

MetricWhat it tells youHow to collect itHealthy signal
Hours saved per person per weekWhether the tools reached real workSelf-reported, weekly, one questionMost people can name a number and a task it came from
Output changeWhether saved time went anywhere usefulBaseline one number per role before the session, measure again afterThe role-relevant number moved, in volume or in quality
Weekly active shareWhether adoption is happening at allOne question on the same formA clear majority, and rising rather than falling week to week

Weekly active share is the leading indicator, and it is the one to watch first. If most of the team used AI for something this week, hours and output will follow on their own. If they did not, no amount of measuring the other two will help — you have a support problem, and it is fixable, but not by collecting more data about it.

On self-reported hours

People will estimate badly. That is fine, and it is not a reason to build something more rigorous. You are not auditing; you are looking for direction and for outliers. Someone reporting zero every week is the most valuable data point on the form, because that is a person you can still help.

The Friday form, in full

Send this every Friday for the first four weeks. It takes under two minutes to complete, and the act of asking does as much work as the answers — a weekly question keeps AI on people's radar during exactly the weeks when the habit is deciding whether to form.

1. Did you use AI for any work task this week? (Yes / No) 2. Roughly how many hours did it save you? (Your best guess is fine — nobody is checking) 3. What did you use it for? Name one specific task. 4. Was there anything you wanted to use it for but couldn't work out how? 5. Anything it got wrong that cost you time?

Questions four and five are the ones people skip when they build this themselves, and they are where the value is. Question four tells you what to teach next. Question five surfaces the bad experience that would otherwise quietly turn one person into a sceptic who tells three colleagues it does not work.

How to read the numbers you get back

  • Most people report a number and can name the task it came from. This is working. Leave it alone and keep asking.
  • Hours are reported but tasks are vague. People are guessing to please you. Push on question three specifically — if nobody can name the task, the hours are not real.
  • A confident majority and a silent minority. The most common pattern, and the most fixable. Go directly to the people reporting nothing. They are almost never resistant; they usually did not know where to start and were too embarrassed to say so in a group.
  • Question four keeps naming the same blocked task. You have found your next session, and it will be the easiest one you ever run.
  • Weekly active share falls in week three. This is the cliff below. Act on it that week, not at the end of the month.

Turning it into a value figure — carefully

You can convert recovered hours into a monetary figure, and finance will eventually ask you to. Do it, but hold the result loosely, because the arithmetic has a flattering bias built in: multiply any plausible number of hours by any plausible cost of an hour across any reasonably sized team, and the annual figure comes out enormous. It comes out enormous whether the programme worked or not.

Two corrections make the number honest. First, recovered hours only become value if they went somewhere — that is what output change measures, and without it you are counting time that may simply have been absorbed. Second, use your own measured hours rather than a benchmark from an article. Treat any published figure, including anything you read here, as a hypothesis to test against your own team rather than a forecast to plan around.

The version of this that survives contact with a CFO is narrow and specific: this team, this task, this many hours, measured over these four weeks, and here is what we did with the capacity. That is a much smaller claim than the ones on most vendors' slides, and it is the only kind that holds up.

The thirty-day cliff

There is a pattern I have now seen in enough organisations to treat as the default expectation rather than a risk: enthusiasm is high for about two weeks, then a normal busy week arrives, and people fall back on the way they know how to do the work. Not because they decided AI was not useful — because the old way is automatic and the new way still requires a decision.

That is why the measurement matters more than it looks. The Friday form is not really reporting. It is the thing that keeps the decision visible during the weeks when it would otherwise quietly stop being made. Teams that keep asking through the dip usually come out the other side with a habit. Teams that measure once at the end find out they lost it, months after they could have done anything about it.

So if you take one thing from this: do not save the measurement until the end. The point of measuring in the first month is not to grade the workshop. It is to catch the fade while it is still catchable.

Every Deployed Kickstart is a half-day session built around what your team actually does, so there is something concrete to measure the following Friday. The Partner programme exists specifically to cover the weeks described above — it is where the fade gets caught.

Frequently asked questions

How do you measure the ROI of an AI workshop?

Measure three things in the first month: hours saved per person per week, output change in a role-relevant number you baselined beforehand, and what share of the team used AI at all this week. The third predicts the other two. Collect all of it with one short weekly form rather than building a reporting apparatus nobody maintains.

Why are workshop satisfaction scores misleading?

Because nearly every feedback form comes back positive — people enjoy a well-run day away from their inbox and will say so. Enjoyment and adoption are barely correlated: sessions with glowing feedback often change nothing, and sessions people found hard going sometimes shift how a whole department works. The only honest test is behavioural — what are people doing four weeks later that they were not doing before?

What should a weekly AI adoption check-in ask?

Five questions: did you use AI for work this week, roughly how many hours did it save, what specific task did you use it for, was there anything you wanted to use it for but could not work out how, and did anything it got wrong cost you time. The last two matter most — one tells you what to teach next, the other surfaces the bad experience that would otherwise turn someone into a quiet sceptic.

Are self-reported time savings reliable enough to use?

They are approximate, and that is acceptable. You are not auditing, you are looking for direction and outliers. The most valuable answer on the form is someone reporting zero every week, because that is a person you can still help. If people report hours but cannot name the task those hours came from, the hours are not real — push on the task question.

What is the thirty-day cliff in AI adoption?

Enthusiasm stays high for around two weeks, then a normal busy week arrives and people fall back on the way they already know how to do the work — not because they decided AI was not useful, but because the old way is automatic and the new way still requires a decision. Teams that keep measuring through that dip usually come out with a habit; teams that measure once at the end find out they lost it months too late to act.

Can you turn recovered hours into a monetary figure?

You can, but hold it loosely — the arithmetic has a flattering bias, because any plausible hours multiplied by any plausible hourly cost across a reasonable team produces an enormous annual figure whether the programme worked or not. Two corrections make it honest: only count hours that demonstrably went somewhere, and use your own measured numbers rather than a published benchmark.

Found this useful? Send it to someone who needs it.

Put this to work

Reading about it is the easy part.

The Deployed Kickstart gets your whole team hands-on with Claude in a single day, mapped to the work you actually do. Tell us where you are and we'll come back within 24 hours.