Skip to content
AI Metric

Chris M.

From minutes saved to value realised

A demonstration that produces a report in five minutes is usually honest. It is also usually answering a different question from the one you are being asked to buy.

What does "time saved" actually mean?

At least six different things, and they are not interchangeable. A percentage that does not say which one it means cannot be checked by anybody.

The number saysWhat it countsWhat it does not tell you
Touch timeMinutes a person is actively workingWhether the report left any sooner
Elapsed timeInstruction to approved issueWhether anyone worked fewer hours
ThroughputAcceptable work finished per hourWhether hours or cost fell
Quality divided by timeGraded output against time usedHow much faster anything got
Released capacityHours freed for something elseWhat the hours were then used for

What does the controlled evidence really show?

TIME PER TASKWORK PER HOURQUALITY DIVIDED BY TIMESLOWER, OR NOTHINGMEASURED EFFECT, BETTER TO THE RIGHT-40%-20%0%20%40%60%80%100%120%140%NO CHANGEExperienced developers, 16 peopletime to finish a task19% SLOWERClinicians, scribe A of twotime spent writing notesNO MEASURABLE CHANGEClinicians, scribe B of twotime spent writing notes9.5% LESSSupport agents, 5,172 peopleissues resolved per hour15% MOREConsultants, 758 peopletime to finish a task25% FASTERDevelopers, 4,867 peopletasks completed26% MOREProfessional writing, 453 peopletime to finish a task40% FASTERLegal tasks, 127 peoplegraded quality divided by time34% TO 140%HOW TO READ THISEach bar is one controlled study. Further right is a better result for the people in it.The colours are four different measurements, not four sizes of the same one, so the barscannot be averaged and no average is drawn. The purple bar is graded quality divided bytime used, which is why it runs past 100 per cent while the time savings behind it were 20 to 28.
Eight controlled results on one axis, better to the right. The colours are four different measurements, not four sizes of the same one, which is why no average is drawn across them. A bounded task can get materially faster, and the same class of tool can also do nothing at all or make an expert slower.Source: Science 2023; Organization Science 2026; QJE 2025; Management Science 2026; NEJM AI 2025; METR 2025; Journal of Law and Empirical Analysis 2026

Bounded professional tasks commonly get 10 to 40 per cent faster. In Science, 453 professionals wrote 40 per cent faster at 18 per cent higher graded quality. In Organization Science, 758 consultants worked about 25 per cent faster inside the model's capability, and were 19 percentage points less likely to be right on the one task outside it. Elsewhere, 5,172 support agents resolved 15 per cent more issues an hour and 4,867 developers completed 26 per cent more tasks.

Then the other half of the same literature. Sixteen experienced developers took 19 per cent longer on code they knew well, while believing they had been 20 per cent faster. In one randomised trial of two ambient clinical scribes, one cut documentation time by 9.5 per cent and the other did nothing measurable. Same technology, same year, opposite results, which is roughly how construction AI actually fails as well.

Very large numbers are not automatically marketing. A 2026 legal trial reported gains of 34 to 140 per cent because productivity there means graded quality divided by time used. The time reductions underneath were 20 to 28 per cent. That is a legitimate measure and a different one.

Why does a task saving shrink across a whole job?

AN ILLUSTRATIVE CHAIN, NOT A MEASURED BASELINEEvery multiplier below is a number you can measure on your own reports in a fortnight.ONE REPORT6.00 hBaseline touch time, firstinstruction to approved issue.TIMES ELIGIBLE SHARE1.50 hOnly a quarter of that isdrafting, the stage the toolreaches.TIMES MEASURED SAVING0.60 hA 40 per cent cut on those stages,at the top of the controlledrange.TIMES ADOPTION0.42 hUsed on 70 per cent of the reportsthat qualify for it.LESS CHECKING0.22 hTwelve more minutes reading thedraft against the evidence.0.22 hours of 6.00. A tool that genuinely cuts drafting time by 40 per cent moved the job by 3.7 per cent.Raise the eligible share from a quarter to three fifths and the same tool nets 15 per cent. Your baseline decides the answer.HOW TO READ THISRead top to bottom. Each bar is what survives the multiplier named to its left.The grey track behind every bar is the whole report, so the shrinking is to scale.Net hours saved equals baseline times eligible share times measured saving times adoption,less verification, exceptions and the cost of running the system.
An illustrative chain, not a measured construction baseline. A tool that genuinely cuts drafting time by 40 per cent moves this report by 3.7 per cent, because drafting is a quarter of the work, the workflow is used on 70 per cent of eligible reports, and checking went up. Change the eligible share and the same tool lands anywhere between 4 and 15 per cent.

Because a survey, inspection or progress report is not prose. It is capture, evidence selection, interpretation, drafting, checking, communication and accountable judgement, and a tool reaches only some of that. Net hours saved is baseline times eligible share times measured saving times adoption, less verification, exceptions and the cost of running the system.

Every one of those is a number you can measure on your own reports in a fortnight. None of them is a number a supplier can know about your firm in advance.

I found no end-to-end randomised trial of construction or surveying report production. Until there is one, 5 to 10 per cent across a whole role is a planning hypothesis to test, not a benchmark to quote.

Where should the automation stop?

ONE REPORT, LEFT TO RIGHTThe test is not how clever the model is. It is whether the output can be checked against something.MACHINE DOES ITYou verify against a sourceDictation intocontrolled fieldsStanding data anddocument metadataPhotograph naming andcross referencingCompleteness andconsistency checksMACHINE DRAFTSA named person owns itDescriptive first draftfrom supplied evidenceSummary of records,cited to the passageConflicting evidencesurfaced, not hiddenPERSON DECIDESEvery timeCause, significance andpriorityContractual position,liability, safetyFinal recommendation andsign offHOW TO READ THISLeft: the output can be compared with the thing it came from, so a machine can produce it.Middle: the draft is evidence linked and a named reviewer is accountable for what it says.Right: the answer depends on judgement, so a person decides and records why.Mandatory for RICS members since 9 March 2026 where AI materially affects a service: a recordedreliability decision by a named surveyor, an AI register, and sampling of high volume automation.
The reporting workflow banded by who decides. The test is not how capable the model is. It is whether the output can be compared with the thing it came from, which is a property of the task rather than of the tool.

Photograph handling, standing data, dictation into controlled fields and completeness checks can be compared against a source, so a machine can do them. Evidence-linked drafting sits in the middle, where model output is a draft and never a source. Cause, significance, contractual position and safety stay with a person, which is the human in the loop matrix doing its job.

The RICS standard on responsible AI use has been mandatory since 9 March 2026 wherever AI materially affects a service. It requires a recorded reliability decision by a named surveyor, an AI register and sampling of high-volume automation. That is a running cost. It is also the control system that makes the saving defensible when somebody challenges the report.

For scale, 13 per cent of UK construction businesses with ten or more employees used any AI technology in June 2026, against 35 per cent across all sectors.

What should you ask before you sign anything?

Demand the denominator. Which outcome, over what boundary, at what quality bar, on whose reports, with adoption and correction time counted in.

"Across 42 matched condition reports in a 90-day pilot, median touch time from evidence pack to approved issue fell 18 per cent after corrections, with no measured quality loss, on 64 per cent of eligible reports" is a claim you can audit.

"Eighty per cent faster" is not a claim at all.

AI Metric is a construction-native AI consultancy. If your team is spending more time operating software than doing their job, book a 30 minute call.