Skip to content
AI Metric

Chris M.

From minutes saved to value realised

Sixteen experienced developers were given AI tools and set to work on code they knew well. They took 19 per cent longer. Asked afterwards, they said it had made them about 20 per cent faster.

That trial is real, recent and randomised. So are the ones showing large gains in the opposite direction. None of them tells you what automating your reports would save. This piece is about the gap between a convincing demonstration and a business case that survives a whole job.

What does "time saved" actually mean?

At least six different things, and they are not interchangeable. A percentage that does not name one cannot be checked.

The number saysWhat it countsWhat it does not tell you
Touch timeMinutes a person is actively workingWhether the report left any sooner
Elapsed timeInstruction to approved issueWhether anyone worked fewer hours
ThroughputAcceptable work finished per hourWhether hours or cost fell
Quality divided by timeGraded output against time usedHow much faster anything got
Released capacityHours freed upWhat those hours were then used for

What does the controlled evidence really show?

TIME PER TASKWORK PER HOURQUALITY DIVIDED BY TIMESLOWER, OR NOTHINGMEASURED EFFECT, BETTER TO THE RIGHT-40%-20%0%20%40%60%80%100%120%140%NO CHANGEExperienced developers, 16 peopletime to finish a task19% SLOWERClinicians, scribe A of twotime spent writing notesNO MEASURABLE CHANGEClinicians, scribe B of twotime spent writing notes9.5% LESSSupport agents, 5,172 peopleissues resolved per hour15% MOREConsultants, 758 peopletime to finish a task25% FASTERDevelopers, 4,867 peopletasks completed26% MOREProfessional writing, 453 peopletime to finish a task40% FASTERLegal tasks, 127 peoplegraded quality divided by time34% TO 140%HOW TO READ THISEach bar is one controlled study. Further right is a better result for the people in it.The colours are four different measurements, not four sizes of the same one, so the barscannot be averaged and no average is drawn. The purple bar is graded quality divided bytime used, which is why it runs past 100 per cent while the time savings behind it were 20 to 28.
Eight controlled results on one axis, better to the right. The colours are four different measurements, not four sizes of the same one, which is why no average is drawn across them. A bounded task can get materially faster, and the same class of tool can also do nothing at all or make an expert slower.Source: Science 2023; Organization Science 2026; QJE 2025; Management Science 2026; NEJM AI 2025; METR 2025; Journal of Law and Empirical Analysis 2026

Bounded tasks do get materially faster. In Science, 453 professionals wrote 40 per cent faster at 18 per cent higher graded quality. In Organization Science, 758 consultants worked about 25 per cent faster inside the model's capability, and were 19 percentage points less likely to be right on the one task outside it. 5,172 support agents resolved 15 per cent more issues an hour, and 4,867 developers completed 26 per cent more tasks.

Then the other half of the same literature, which includes the developer trial this piece opened with. In a randomised trial of two ambient clinical scribes, one cut documentation time by 9.5 per cent and the other did nothing measurable, a pattern construction AI keeps repeating.

Very large numbers are not automatically marketing. A 2026 legal trial reported gains of 34 to 140 per cent because productivity there means graded quality divided by time used. The time reductions underneath were 20 to 28 per cent. A legitimate measure, and a different one.

Why does a task saving shrink across a whole job?

AN ILLUSTRATIVE CHAIN, NOT A MEASURED BASELINEEvery multiplier below is a number you can measure on your own reports in a fortnight.ONE REPORT6.00 hBaseline touch time, firstinstruction to approved issue.TIMES ELIGIBLE SHARE1.50 hOnly a quarter of that isdrafting, the stage the toolreaches.TIMES MEASURED SAVING0.60 hA 40 per cent cut on those stages,at the top of the controlledrange.TIMES ADOPTION0.42 hUsed on 70 per cent of the reportsthat qualify for it.LESS CHECKING0.22 hTwelve more minutes reading thedraft against the evidence.0.22 hours of 6.00. A tool that genuinely cuts drafting time by 40 per cent moved the job by 3.7 per cent.Raise the eligible share from a quarter to three fifths and the same tool nets 15 per cent. Your baseline decides the answer.HOW TO READ THISRead top to bottom. Each bar is what survives the multiplier named to its left.The grey track behind every bar is the whole report, so the shrinking is to scale.Net hours saved equals baseline times eligible share times measured saving times adoption,less verification, exceptions and the cost of running the system.
An illustrative chain, not a measured construction baseline. A tool that genuinely cuts drafting time by 40 per cent moves this report by 3.7 per cent, because drafting is a quarter of the work, the workflow is used on 70 per cent of eligible reports, and checking went up. Change the eligible share and the same tool lands anywhere between 4 and 15 per cent.

Because a report is not prose. It is capture, evidence selection, interpretation, drafting, checking, communication and accountable judgement, and a tool reaches only some of that. Net hours saved is baseline times eligible share times measured saving times adoption, less verification, exceptions and the cost of running the system. Every one is a number you can measure on your own reports in a fortnight, and none is one a supplier can know in advance.

I found no end-to-end randomised trial of construction or surveying report production. Until there is one, 5 to 10 per cent across a whole role is a planning hypothesis to test, not a benchmark to quote.

Where should the automation stop?

ONE REPORT, LEFT TO RIGHTThe test is not how clever the model is. It is whether the output can be checked against something.MACHINE DOES ITYou verify against a sourceDictation intocontrolled fieldsStanding data anddocument metadataPhotograph naming andcross referencingCompleteness andconsistency checksMACHINE DRAFTSA named person owns itDescriptive first draftfrom supplied evidenceSummary of records,cited to the passageConflicting evidencesurfaced, not hiddenPERSON DECIDESEvery timeCause, significance andpriorityContractual position,liability, safetyFinal recommendation andsign offHOW TO READ THISLeft: the output can be compared with the thing it came from, so a machine can produce it.Middle: the draft is evidence linked and a named reviewer is accountable for what it says.Right: the answer depends on judgement, so a person decides and records why.Mandatory for RICS members since 9 March 2026 where AI materially affects a service: a recordedreliability decision by a named surveyor, an AI register, and sampling of high volume automation.
The reporting workflow banded by who decides. The test is not how capable the model is. It is whether the output can be compared with the thing it came from, which is a property of the task rather than of the tool.

Photographs, standing data, dictation and completeness checks can be compared against a source, so a machine can do them. Evidence-linked drafting sits in the middle, where model output is a draft and never a source. Cause, contractual position and safety stay with a person, per the human in the loop matrix.

The RICS standard on responsible AI use has been mandatory since 9 March 2026 wherever AI materially affects a service: a recorded reliability decision by a named surveyor, an AI register and sampling of high-volume automation. A running cost, and also the control system that makes the saving defensible when the report is challenged.

For scale: 13 per cent of UK construction firms with ten or more staff used any AI in June 2026, against 35 per cent across all sectors.

What should you ask before you sign anything?

Demand the denominator. Which outcome, over what boundary, at what quality bar, on whose reports, with correction time counted in.

"Across 42 matched condition reports in a 90-day pilot, median touch time from evidence pack to approved issue fell 18 per cent after corrections, with no measured quality loss, on 64 per cent of eligible reports" is a claim you can audit.

"Eighty per cent faster" is not a claim at all.

AI Metric is a construction-native AI consultancy. If your team is spending more time operating software than doing their job, book a 30 minute call.