From minutes saved to value realised
Sixteen experienced developers were given AI tools and set to work on code they knew well. They took 19 per cent longer. Asked afterwards, they said it had made them about 20 per cent faster.
That trial is real, recent and randomised. So are the ones showing large gains in the opposite direction. None of them tells you what automating your reports would save. This piece is about the gap between a convincing demonstration and a business case that survives a whole job.
What does "time saved" actually mean?
At least six different things, and they are not interchangeable. A percentage that does not name one cannot be checked.
| The number says | What it counts | What it does not tell you |
|---|---|---|
| Touch time | Minutes a person is actively working | Whether the report left any sooner |
| Elapsed time | Instruction to approved issue | Whether anyone worked fewer hours |
| Throughput | Acceptable work finished per hour | Whether hours or cost fell |
| Quality divided by time | Graded output against time used | How much faster anything got |
| Released capacity | Hours freed up | What those hours were then used for |
What does the controlled evidence really show?
Bounded tasks do get materially faster. In Science, 453 professionals wrote 40 per cent faster at 18 per cent higher graded quality. In Organization Science, 758 consultants worked about 25 per cent faster inside the model's capability, and were 19 percentage points less likely to be right on the one task outside it. 5,172 support agents resolved 15 per cent more issues an hour, and 4,867 developers completed 26 per cent more tasks.
Then the other half of the same literature, which includes the developer trial this piece opened with. In a randomised trial of two ambient clinical scribes, one cut documentation time by 9.5 per cent and the other did nothing measurable, a pattern construction AI keeps repeating.
Very large numbers are not automatically marketing. A 2026 legal trial reported gains of 34 to 140 per cent because productivity there means graded quality divided by time used. The time reductions underneath were 20 to 28 per cent. A legitimate measure, and a different one.
Why does a task saving shrink across a whole job?
Because a report is not prose. It is capture, evidence selection, interpretation, drafting, checking, communication and accountable judgement, and a tool reaches only some of that. Net hours saved is baseline times eligible share times measured saving times adoption, less verification, exceptions and the cost of running the system. Every one is a number you can measure on your own reports in a fortnight, and none is one a supplier can know in advance.
I found no end-to-end randomised trial of construction or surveying report production. Until there is one, 5 to 10 per cent across a whole role is a planning hypothesis to test, not a benchmark to quote.
Where should the automation stop?
Photographs, standing data, dictation and completeness checks can be compared against a source, so a machine can do them. Evidence-linked drafting sits in the middle, where model output is a draft and never a source. Cause, contractual position and safety stay with a person, per the human in the loop matrix.
The RICS standard on responsible AI use has been mandatory since 9 March 2026 wherever AI materially affects a service: a recorded reliability decision by a named surveyor, an AI register and sampling of high-volume automation. A running cost, and also the control system that makes the saving defensible when the report is challenged.
For scale: 13 per cent of UK construction firms with ten or more staff used any AI in June 2026, against 35 per cent across all sectors.
What should you ask before you sign anything?
Demand the denominator. Which outcome, over what boundary, at what quality bar, on whose reports, with correction time counted in.
"Across 42 matched condition reports in a 90-day pilot, median touch time from evidence pack to approved issue fell 18 per cent after corrections, with no measured quality loss, on 64 per cent of eligible reports" is a claim you can audit.
"Eighty per cent faster" is not a claim at all.