AI Tools Academy
0 / 117 (0%)

AI for managers

Is AI actually saving your team time?

About 14 minutesPractises: Task judgement, Verification, Human responsibility

Written by AI Tools AcademyLast checked on 27 September 2026

Helpful first: Which tasks are suitable for AI?

Practice files for this page

AI time measurement template (CSV spreadsheet)

Fictional practice data. No real people or organisations.

Erin Vaughan, Fernway's Managing Director, asks Priya Shah a straightforward question: "We've had Copilot Chat for three months. Is it saving us time?" Priya realises she doesn't know. People say it helps. Some use it every day. But nobody measured how long anything took before, and nobody has counted the time spent checking what comes out.

This guide shows how to answer Erin's question honestly. It covers which numbers to ignore, what to measure instead, and a simple way to do it without turning your team into a time-and-motion study. It includes a template and two worked examples: one where AI saved real time, and one where it didn't once checking was counted.

The numbers that don't tell you much

Some figures are easy to collect and sound impressive, but say little about whether work is getting better or quicker. These are sometimes called vanity metrics.

Number of licences or active users. Tells you people opened the tool. Not what they did with it, or whether it helped.

Number of prompts or chats. More prompts could mean more useful work, or more attempts to get a usable answer.

Self-reported "hours saved" with no baseline. People are poor at estimating how long a task used to take, and the moment a draft appears feels like the work is nearly done. The checking and fixing that follows is easy to forget.

Time to first draft. AI produces a draft in seconds. That tells you nothing about how long it takes to get to something you can send.

None of these are useless. Usage figures can show whether people have access and are trying the tool. They can't answer whether it's saving time.

What to measure instead

For each task you want to assess, collect a few simple things.

Baseline task time. How long the task takes by hand, from start to finished. Time it over three to five runs and take the average, because tasks vary.

Time with AI, broken into parts. Preparing the material and writing the prompt. Generating the draft. Reviewing and correcting it. Splitting these shows where the time really goes.

Rework. Time spent later fixing something that went out wrong: a correction email, a customer query, a figure re-done. This is the part people most often leave out.

Errors. How many errors were found before sending, and how many after. Errors found after sending are the ones that matter to customers and colleagues.

Quality. Is the result as good as, better than or worse than the hand-made version? A simple 1 to 5 rating from the person who receives the work is enough.

Employee experience. How did it feel to do? A task that saves five minutes but feels stressful and fiddly may not last. One that takes the same time but removes a dreaded chore may be worth keeping.

Customer impact. For customer-facing work, did anything change: complaints, queries, response times, comments?

Frequency. How often the task happens. A 20-minute saving on a daily task is worth far more than the same saving on a yearly one.

A simple way to measure

You don't need software or a project team. For two to four weeks:

  1. Pick one to three tasks. Choose ones that are repeated often and have been rated suitable. The task suitability framework helps you choose.
  2. Time a baseline. Before using AI, time three to five runs of the task done by hand. If AI is already in use, ask someone to do a few runs by hand for comparison, or use a task that is still done both ways.
  3. Time the AI runs honestly. Record preparation, drafting, checking and correction separately. Note any rework later.
  4. Record errors and quality. Ask the person receiving the work to rate it, without telling them which version used AI if you can.
  5. Compare averages, not best cases. One quick run tells you little. Look at the average over several.

Explain to the team why you're measuring. You're assessing the task, nobody's performance, and you'll share the results. People who think they're being watched may rush, or quietly skip checks to look fast, which defeats the point.

Download the AI time measurement template. It opens in Excel or Google Sheets and includes three fictional example rows to show how it works.

Show the columns in the template
  • Row type: a fictional example, or your own data
  • Task, person or role and date
  • Method: by hand or with AI
  • Setup and prompting (minutes): gathering material, removing anything that can't go into the tool, writing the prompt
  • Doing or drafting (minutes): the main work, or waiting for the AI draft
  • Review and correction (minutes): checking against the source and fixing
  • Rework later (minutes): fixing anything found after it was sent
  • Total (minutes): the sum of the four time columns
  • Errors found before sending and errors found after sending
  • Quality, 1 to 5, rated by the recipient
  • How it felt, 1 to 5, rated by the person doing it
  • Customer impact notes
  • Times per month
  • Notes

Worked example 1: AI saved real time

Fictional example. Leah Bennett writes the weekly operations update from her notes and the delivery tracker.

Baseline by hand. Leah timed four weeks: 48, 55, 47 and 50 minutes. The total is 200 minutes, so the average is 50 minutes a week.

With AI. For the next four weeks she drafted in Copilot Chat from her notes, then checked every date, supplier name and figure against the tracker. Her totals were 27, 24, 26 and 23 minutes. The total is 100 minutes, so the average is 25 minutes a week.

A typical AI week broke down like this:

Part of the taskMinutes
Tidying notes and writing the prompt10
Generating the draft2
Checking against the tracker and correcting13
Rework after sending0
Total25

Errors. Checking caught one wrong delivery date in week two. Nothing needed correcting after sending in either period.

Quality. Priya rated the updates about the same in both periods. The team said the AI versions were easier to scan.

The saving. 50 minus 25 is 25 minutes a week. Allowing 46 working weeks a year for leave and bank holidays, that's 25 × 46 = 1,150 minutes, or about 19 hours a year for one weekly task.

This is a real saving because Leah's notes contain all the facts, the output goes to an internal audience, and checking against the tracker is quick. She still spends more time checking than generating, which is how it should be.

Worked example 2: AI didn't save time once checking was counted

Fictional example. Grace Lynch prepares a monthly summary of supplier statements for Dan Okafor, flagging any line that doesn't match Fernway's records.

Baseline by hand. Grace timed three months: 38, 42 and 40 minutes. The total is 120 minutes, so the average is 40 minutes a month.

With AI. She tried pasting the statement lines (with bank details removed) into the approved tool and asking it to summarise and flag mismatches.

Part of the taskMinutes
Preparing the data and removing bank details5
Generating the summary1
Checking every line against the ledger30
Rework: correcting a query sent to a supplier about a figure the AI had misread10
Total46

Errors. Checking caught two totals that were wrong in the AI summary. One more error was only found after Grace had emailed a supplier about it.

The result. 46 minutes with AI against 40 by hand: 6 minutes a month slower. Over 12 months, that's 6 × 12 = 72 minutes a year lost, plus an awkward email to a supplier.

The draft appeared in about six minutes, which felt like a big saving on 40. It wasn't. The task is exact matching of figures, the AI's errors looked plausible, and the only safe way to check was to go through every line, which is the task itself.

Grace switched to a spreadsheet lookup formula for the matching, and used AI only to help her write the formula. That's a better use of the tool for this job.

Using the time-saved calculator

Before you measure anything, the free time-saved calculator gives a rough, deliberately cautious estimate of where the biggest opportunities might be. It asks how often a task happens, how long it takes and how much of it AI could help with, then halves the assisted time as an allowance for checking.

Use it to decide which tasks are worth measuring. Then measure them properly with the template. The calculator gives you a guess, and the measurement tells you what actually happened.

Reporting back honestly

When Priya reported to Erin, she didn't give a single "hours saved" figure. She reported task by task:

  • Weekly operations update: about 25 minutes a week saved, no drop in quality, keep going.
  • Supplier statement summary: slower with AI once checking was included, switched to a spreadsheet formula instead.
  • Standard supplier chaser emails: saving a few minutes each, but still measuring.

That report is more useful than a headline number. It shows where AI helps, where it doesn't, and why. It also shows the team that honest results, including "this didn't work", are welcome.

Turn your measurements into a short reportAny approved work AI tool
Below are my team's timings for tasks done by hand and with AI. Using only these figures, write a short report for a senior manager. For each task: the average time by hand, the average time with AI including checking and rework, the difference per run and per year using the frequency given, and any errors found after sending. Show the calculation for each figure so I can check it. Say plainly where AI did not save time. Do not add any figures that are not in the data. UK English, under 300 words.

[paste your measurement rows, with names removed]

Why this works: It keeps the AI to your figures, asks it to show the working so you can check the arithmetic, and asks for the tasks that didn't work as well as those that did.

Check every figure in the report against your spreadsheet before you send it. AI tools can make arithmetic mistakes, and a report about time savings with wrong sums in it undermines itself.

Common mistakes

Measuring only the drafting. The draft is the quick part. Checking and rework are where the time goes.

Comparing a good day with AI to a bad day without. Use averages over several runs.

Ignoring rework. An error found after sending costs far more than one found before.

Measuring people, not tasks. If staff feel their speed is being judged, they'll rush and skip checks. Make it clear you're assessing the task.

Stopping at the first result. People get quicker with practice, and tools change. Re-measure after a couple of months.

Assuming a saving means fewer people are needed. Time saved on one task usually goes into other work. What happens to it is a decision for the organisation, and the introducing AI without panic guide covers how to talk about it honestly.

Rolling this out to a team?

We run practical, remote training for teams on safe, useful AI at work, using fictional practice data so nobody has to use real information to learn.

Discuss team training