Measure automation in time, not money
Every pounds-saved claim about automation starts an argument. Say a system "saves £30,000 a year" and a sensible director immediately attacks the arithmetic: which charge-out rate, whose overhead, and saved compared to what? Would that work really have been done, or paid for, in the world without the system? The number collapses under questioning because it was never a measurement; it was a rate multiplied by an estimate multiplied by a counterfactual.
Hours do not have this problem. Hours are observable. Either the estimator spent Tuesday evening retyping enquiry details or she did not. Either the quote went out in two days or in nine. You can stand in the office and watch the difference. So measure automation in time returned and cycle times, publish those numbers, and let every reader multiply by whatever rate they believe. The multiplication is theirs; the measurement is yours.
Why do money claims fail where time claims survive?
Because a money claim smuggles in three assumptions, and a sceptic only needs to reject one.
The rate: is an hour of a QS worth their salary, their charge-out, or their opportunity cost? The overhead: does the saved hour shed any actual cost, or does the salary get paid regardless? The counterfactual: would the firm otherwise have hired, paid overtime, or simply not done the work? Reasonable people disagree on all three, so any pounds figure is contestable by construction (in both senses).
A time claim asserts only what happened: this task took four hours a week before; it takes twenty minutes now. The reader who wants pounds can apply their own assumptions, and the ones they choose they will not argue with. This is the same honesty rule that applies everywhere on this subject: as with the cost of doing nothing, the moment you invent a precise-sounding number, you hand the sceptic the whole argument.
What should you actually measure?
Three families of number, all countable without a consultant.
First, hours per week returned, per person, per task. Not "the team saves time": the diary took Dave five hours a week, it now takes forty minutes of review. Second, response times: how long an enquiry waits for an acknowledgement, a query for an answer. Third, cycle times: enquiry to quote, completion to invoice, month end to accounts. Cycle times matter because they capture value that never appears in anyone's timesheet; a quote that goes out in two days instead of nine wins work that the nine-day version never sees, which is one reason automation that runs without being asked shows up in cycle times first.
National productivity statistics, such as the output measures in the ONS construction data, track an industry-level version of the same idea: output against hours worked. The Construction Leadership Council has had productivity on its agenda for years. You are simply running the firm-sized version, where the hours are ones you can actually see.
What does a before and after measurement plan look like?
Small, specific and written down before anything is built.
| Measure | How to baseline it | After automation |
|---|---|---|
| Hours per week on the task, per person | Two ordinary weeks of honest notes from the people who do it | Same notes, same people, same fortnight length |
| Enquiry response time | Timestamps already in the inbox: received versus first reply, last 20 enquiries | Same timestamps, next 20 enquiries |
| Enquiry to quote cycle | Dates on the last 20 quotes against their enquiry dates | Same count, measured the same way |
| Completion to invoice | Job completion dates against invoice dates, last quarter | Next quarter, same definition of completion |
| Chasing effort | Count the chase emails and calls in a normal week | Count again; the system logs its own |
Nothing in that table needs new software to measure. Most of it is timestamps you already have and a fortnight of honest notes.
Why must the baseline come before the automation?
Because after the system exists, the before disappears, and with it your only defensible number.
This is the discipline almost everyone skips. The project gets approved, the build starts, and six months later somebody asks what it achieved. Now the only available answer is a reconstruction: people guessing what the diary used to take, remembering the bad weeks and forgetting the ordinary ones. A reconstructed baseline is exactly as contestable as a pounds claim, and it deserves to be.
Baselining first costs two weeks of light effort and zero delay, because it runs while the automation is being scoped. It also improves the scope: the act of timing a process usually reveals that the real bottleneck is not where everyone assumed. A pilot with a pre-agreed baseline and a pre-agreed success number is the core of what a good AI pilot looks like; without one, a pilot is just a demo with a start date.
One warning from experience: measure tasks, not people. The moment staff believe the stopwatch is aimed at them rather than at the process, the notes turn fiction. Say plainly that the question is "where do your hours go?" and never "are you fast enough?", and mean it.
How do you report it without spin?
State the observation, show the method, and do the reader's multiplication only as an open example.
"The diary took five hours a week of supervisor time; it now takes forty minutes of review. At a fully costed £40 an hour that would be roughly £9,000 a year, but apply your own rate." The first sentence is a measurement nobody can argue with. The second is arithmetic anyone can check and adjust. AI Metric baselines every workflow automation engagement this way, mostly out of self-interest: a client who watched the baseline being taken never disputes the result.
Time is the honest currency of automation. Count it before, count it after, and let the pounds argue for themselves.