Stop measuring AI by time saved
"It saved me two hours" is a useful sentence and a terrible strategy. Those two hours can become a shorter day, more site presence, more meetings, a wider span of control, or a more thorough piece of work. A basic return calculation records all five identically.
What were the answers to the previous five questions?
1. Why is time saved a poor primary measure of AI value in construction?
Because saving two hours is not the same as creating two hours of value. The same saving can produce a shorter day, more site presence, more meetings, a wider span of control or a more thorough piece of work. Those are five different outcomes and a basic calculation records all of them as two hours.
2. A report that took four hours now takes forty minutes and changes no decision. What has been achieved?
Waste has been made more efficient. If nobody does anything differently because the report exists, the right intervention is not producing it faster, it is asking why it exists.
3. What are the six categories worth measuring instead, and what does each one ask?
Time: what work disappeared, not what ran faster. Quality: did errors, omissions and rework fall. Risk: were issues found earlier and decisions better evidenced. Commercial: was leakage prevented and entitlement evidenced. Capacity: what were the released hours deliberately spent on. Human capability: did the professional get better, or only faster.
4. Why must a baseline be captured before anyone knows a tool is coming?
Because measurement changes behaviour. Once people know the process is being observed they work differently, and the comparison stops being like for like. Baseline first, announcement second.
5. What are the three honest outcomes of a six week measurement cycle?
Expand, because the evidence shows value. Adjust, because there is value and the implementation is wrong. Abandon, because there is not. A cycle that can only conclude expand was never a measurement exercise.
What is wrong with the standard business case?
It multiplies. Two hours saved three times a week per project manager, thirty project managers, forty eight weeks, converted to salary cost, and the slide says the software saves hundreds of thousands of pounds a year.
Nobody's salary disappears. Nobody works a shorter week. The project does not finish sooner and the company does not automatically win another job. The arithmetic is correct and the conclusion is unsupported, and construction has been through this with document management, mobile applications, dashboards and ERP already.
The national data says the same thing more politely. Department for Science, Innovation and Technology research found 56 per cent of businesses already using AI reporting increased employee productivity, while 77 per cent had seen no change in revenue and 12 per cent reported an increase (DSIT). Productivity gains are real and they are not converting themselves. Somebody has to decide what the capacity is for.
What should be measured instead?
Six categories, of which time is only the first and the least interesting. Between them they describe what actually changed in the business rather than how long something took.
| Category | A measure that works | A measure that does not |
|---|---|---|
| Time | Manual information transfers removed from the process | Estimated hours saved per user |
| Quality | Rework rate on the deliverable, before and after | Perceived improvement |
| Risk | Interval between a problem being created and being recognised | Number of documents processed |
| Commercial | Unrecorded change identified, late notices detected, unsupported cost challenged | Value of the contract the tool touches |
| Capacity | What released hours were deliberately reallocated to | Hours released |
| Human capability | Frequency with which staff challenge an output, and source verification rate | Training completions |
| Cost of ownership | Licences, integration, the knowledge layer build and its upkeep, added review time, assurance and governance effort, measured over the same weeks | The licence cost |
The last row is the one most often left out, and leaving it out is the same error the part opens by attacking. The third row is the one worth arguing for internally. Construction losses do not originate when the financial consequence appears. They originate at an unresolved design interface, an incomplete instruction, a missing approval, a subcontractor falling behind. The problem exists; the business has not recognised it yet. If a commercial risk previously surfaced in the month nine CVR and now surfaces in month three, the value is not the twenty minutes saved producing the CVR. It is six months of ability to do something about it.
How do you actually run the measurement?
Pick one recurring process. Not five. One that is well defined, has clear inputs and outputs, and happens often enough to give several data points: weekly progress reporting, monthly commercial reconciliation, subcontractor invoice review, site diary production.
In a contracting business the benefit lands in three places and nowhere else. Margin: fewer write-downs at valuation and fewer late reversals. Cash: earlier and better evidenced applications, fewer pay less notices received, fewer contra-charges absorbed. Provisions: a smaller risk provision and lower dispute cost. A claimed benefit that cannot be traced to one of those three has not reached the business.
Keep the recording under two minutes per cycle or it will not happen. Capture elapsed time and actual labour time separately, the number of systems accessed, how many times information was re-keyed, errors discovered, rework required, and what the released time went to. Then compare against all six categories and be honest about the ones that got worse.
A worked example of the shape an honest result takes, illustrative rather than measured and presented as such: a weekly progress report where total effort falls from 5.5 hours to 3, data gathering goes to zero, and review time rises from half an hour to nearly two because people are checking the output properly. Rework falls from 15 per cent to 3 per cent. Two new AI errors per report appear and are caught in review. The number worth tracking is the one nobody has: errors that survived review and were found later, which needs a downstream check and is the only quality measure that tests the part 7 failure modes rather than assuming them away. The headline is a 45 per cent time reduction. The story is that manual transcription disappeared, quality rose, and a new class of error was introduced that requires expertise to catch. Your own numbers will differ, which is the entire argument for taking the baseline.
What should the board be asked instead?
Different questions, which produce different investments.
- Not how many hours can AI save us, but what would become possible if this work disappeared.
- Not how quickly can it produce the report, but does the report lead to an earlier or better decision.
- Not how many documents can it read, but what can it identify that we currently miss.
- Not how much administration can we automate, but what valuable work will our people do instead.
And one uncomfortable one, borrowed from the productivity evidence: what did we deliberately do with the capacity this released? RICS found that roughly one in five UK respondents say their firm never measures productivity at all, and that of the five regions surveyed, UK respondents were the most sceptical about automation as an intervention (RICS, 2026). Those two findings are related. A sector that does not measure cannot be persuaded by evidence, and so is left with vendor arithmetic and instinct.
Construction already understands prevention. We inspect before covering up, review designs before installation, plan lifts before lifting and test materials before accepting them, and we do not value any of it by the minutes it saves. AI assurance deserves the same treatment. See measuring automation in time not money and what a proper AI audit looks like.
One part left. part 13, competence, clients and the first ninety days, published Saturday, 29 August 2026 closes the syllabus with the thing that decides whether any of it holds: whether your people are still competent to challenge the machine, and what you say when a client asks.
What five questions should you be able to answer now?
Attempt these before tomorrow. Each has a defensible answer, and each is answered at the top of the next part.
- What is the actual objective of AI training in a construction business, and how would you know it had been met?
- Why should training deliberately include AI answers that are wrong?
- Which AI uses should be disclosed to a client, and which need not be?
- What are professional indemnity insurers now asking about AI, and what document answers all of it?
- If you did one thing in the next ninety days, what should it be?
Which sources is this part built on?
Every figure quoted above resolves to one of these. Each was checked before publication.