← All guides

How to tell whether AI actually helped your team

Did AI make the job easier? Include checking, corrections and the next person's work when you compare the results.

Alongside People · 5 min read ·

It's a good feeling when a task that usually takes half an hour produces a draft in two minutes. You can see why the team wants to try it again.

Before calling it a time saving, follow that draft a little further. Who checks it? What do they have to fix? Does the next person have what they need to get on with their job? The answers tell you whether the whole piece of work has become easier.

We start an AI trial with a small question the team can answer from experience. Something like: “Can we give delivery a complete customer handover with less chasing?” That gives everyone a reason to try the change and a way to judge it.

Agree on what better would look like

Let's use that handover as a hypothetical example. An account manager gathers the customer's requirements and commitments into a brief. Delivery reads it and starts the work, sometimes coming back with questions.

Sit down with both sides and ask what would make a difference. Perhaps preparation takes too long. Perhaps the same detail is missing every time. Perhaps delivery gets a brief quickly but waits two days for an answer before it can begin.

Choose a few things to follow: time spent preparing and checking, missing details delivery has to chase, and the wait until delivery can start. Define them together. If you count a handover as complete, agree on what needs to be there.

Also decide which mistakes would make the trial unacceptable. An invented customer commitment deserves attention even if the draft arrived quickly. The people responsible for the work should help set those limits.

Look at a few recent jobs first

Take recent handovers and work through what happened. Include ordinary ones and awkward ones. Record the effort involved and any repairs needed later. This gives you a starting point for comparison, often called a baseline.

Be clear about what those examples cover. A handful of easy jobs completed by your most experienced person won't tell you how the process works for everyone. You can still learn from them; keep the limits alongside the result.

During the trial, note other changes too. A clearer intake form, a different mix of customers or an extra pair of hands may affect the outcome. You'll want to remember those details when deciding how much of the improvement to attribute to AI.

Add up the checking as well as the making

Here are some invented numbers to show how the comparison works. They aren't a prediction or a client result.

The old handover takes 30 minutes to prepare and 10 minutes to check. For this example, all 40 minutes require someone's attention and the steps happen one after another. With AI assistance, gathering information takes 5 minutes, generation takes 2, review takes 18 and corrections take 7. If those steps also happen one after another, that's 32 minutes elapsed. If generation runs unattended, the person's active effort is 30 minutes.

The elapsed-time difference is 8 minutes; the active-effort difference is 10. They answer different questions, so keep them separate. Comparing the old preparation time with the generation time alone would suggest a much larger saving and leave most of the new work uncounted.

Use the same definition on both sides of the comparison. Also account for setup, maintenance and sorting out failures across the trial. They may be worthwhile costs; they belong in the picture.

This doesn't need to become an elaborate exercise. A short record for each job can show where time went and help you spot which part needs attention next.

Ask the person who receives the work

Give delivery a say in whether the change helped. Can they start confidently? Do they still have to search for details? Is anything important wrong?

Record the kind of problem as well as how often it happens. A few wording corrections are different from an incorrect promise. It matters whether someone caught the error before acting on it or discovered it after work had begun.

Paul explores this practical side of verification: checking whether the result serves its purpose. A successful technical run tells you the software ran. Delivery can tell you whether the handover helped them deliver.

Talk to the person checking the drafts too. Are errors easy to recognise? Has a dull task become easier, or has review become tiring? Do they know when to stop and ask for help?

Keep what people say alongside what you measure. “I feel less rushed” is useful feedback. A recorded reduction in review time is a different kind of evidence. Both can help you decide what to try next.

Decide what the result lets you do

If the trial releases time, ask what the team can use it for. It might make room for better customer conversations, reduce a backlog or make a busy week more manageable. Minutes released don't automatically become cash savings, and you can't assume every minute will be put to another use.

At the end, write down what changed, what you examined, what improved and what became harder. Include the gaps. A small trial may give you enough confidence to try a few more cases without proving the approach will work everywhere.

You might keep it, narrow the job, improve the inputs or stop. Any of those can be a sound decision when you can explain it. Give someone responsibility for the next step and set a time to look again.

Before your next AI experiment, finish this sentence with the team: “We'll know this helped when...” Then choose a few recent jobs to compare it with. You'll have a much clearer conversation when the results come back.

If you've already tried something and the result is hard to pin down, tell us what you hoped would improve. We can work through the comparison together.

Let’s talk it through

What have you tried?
What’s next?

Tell us what looks promising, what’s getting in the way, or what you’re still trying to work out. A few sentences is plenty to begin.

Let’s talk