Skip to content
AI Metric

Chris M.

The most valuable AI on a project may save nobody any time

Everything in commercial management reduces to one question: what can we prove? A system that assembles evidence continuously changes the answer, and it does not do so by being fast.

What were the answers to the previous five questions?

1. Why is commercial management described as an evidence game rather than a numbers game?

Because every commercial position resolves to what can be proved. A variation turns on the contract terms, the instruction, the correspondence, the programme movement and the cost incurred. The arithmetic is the easy part; assembling and defending the evidence is the job.

2. What is a challenge function, and why might it save no time at all?

A system that tests a professional assessment against the available evidence and reports where they diverge. It often adds twenty minutes of investigation. Its value is in the accuracy of the position that results, not in the duration of producing it.

3. A CVR says a package is 74 per cent complete. What would you compare that against?

Start with remeasure against the contract rules of measurement, because that is the only direct evidence and everything else is proxy. Then materials on and off site separated from installed work, programme progress, inspection records and open non-conformances, site photographs, committed and accrued cost, and previous assessments. Then walk the works. Where those cluster somewhere other than 74, the assessment needs justifying rather than accepting.

4. Why does a contemporaneous commercial record beat a reconstructed one, and what does that mean for final accounts?

Because reconstruction is memory plus whatever survived, and it is visibly weaker under scrutiny. If the record is built as events happen, the final account becomes validation of an existing position rather than an archaeological exercise months after the fact.

5. What should happen when a commercial manager overrides an assessment, and what makes the difference between a strong and a weak position later?

The override is recorded with the specific reasoning and the person taking responsibility. Weak: the number changes and nobody writes down why, so the file contradicts itself under challenge. Strong: the file shows the same evidence assessed differently, by a named person, for stated reasons.

What is a quantity surveyor actually doing all day?

Reconciling. Finding the latest subcontract, reading correspondence, checking applications, updating trackers, comparing forecasts, extracting quantities, preparing CVRs, reviewing variation logs, chasing evidence, reconstructing what happened three months ago.

None of that is unimportant, and almost none of it is why the business employs an experienced surveyor. The commercial position of a project only becomes visible once contracts, applications, progress records, instructions, variations, cost data and programme have all been reconciled, and that reconciliation is exactly the activity a machine is good at.

A typical project carries twenty thousand or more emails, thousands of drawings, hundreds of meeting minutes, site diaries, photographs, subcontract documents, payment records and programme updates. No surveyor reads all of it. They sample, and they rely on experience about where problems usually appear. That works, and it misses things, and the things it misses are not random.

What does a challenge function do that a report does not?

It disagrees with you, with references.

Assessed at 74%Evidence range 61 to 68%Settled at 64% after investigation0%10%20%30%40%50%60%70%80%90%100%ProgrammeInspectionsPhotographsCommitted costApplicationsTime saved: none. The exercise added about twenty minutesand moved the reported position by ten points.
Illustrative. A subcontract package certified at 74 per cent against an evidence range of 61 to 68 per cent drawn from six separate lines of evidence. After investigation the position settles at 64. Over-certification here is cash out of the door against work that does not exist, and unrecoverable if the subcontractor fails.

Run that through a conventional return on investment calculation and it fails: it consumed time and produced no saving. Run it through the question the board actually cares about and it succeeds. On a two point four million pound package, ten points is two hundred and forty thousand pounds of value carried in a reported position the evidence did not support, and the intervention happened in month three rather than month nine. Construction recorded more company insolvencies than any other industry in the twelve months to June 2026, which is the reason over-certification against a subcontractor is not an accounting error.

One carve-out before the table. Where a notice operates as a condition precedent, the system is a safety net and never the detector of record. A missed detection there does not produce a late notice, it produces no entitlement at all, and the recall targets that are acceptable elsewhere are not acceptable here. The contractual diary, the notice register and the named person responsible for them stay the primary control. The statutory payment regime under section 110A of the Construction Act is the same shape: a missed pay less deadline is an immediate and unrecoverable consequence, not a delay.

The same pattern applies across the commercial workload:

ActivityWhat the system contributesWhat stays with the surveyor
Change assessmentThe contract definition, the specification assumption, the site evidence, the gap between them, and what evidence is missingWhether it meets the contractual definition of a change, which valuation basis this contract actually provides, and the position taken
Application reviewLine by line reconciliation against the contract basis, previous applications and programme status, reporting duplicates, unsupported items and exceptions for the surveyor to assessThe notified sum, the payment or pay less notice that carries it, and the basis of calculation stated in that notice
Change detectionContinuous scanning of correspondence and drawings for potential instructed change never registered, citing the source passage and drawing no conclusion on whether an instruction was validly givenWhether an instruction was given, by someone with authority to give it, and what to do about it
Notice monitoringIdentification of events that may start a notice period running, measured against the notice provisions as amended and the date the event arose or ought to have been apparentWhether to issue, in what form, and under whose authority
Payment noticesDue dates, notice deadlines and notified sums tracked across every subcontract, counting down to the final date for payment and the pay less deadlineThe notice itself, its content, its basis of calculation, and issuing it in time
Forecast to completeActual productivity, remaining quantities, burn rate, and the divergence from the stated forecast. A quantity the system produces is a measurement and carries the same professional liability as one you took yourselfThe forecast the business publishes and defends
Final accountA live chronology of instructions, changes, approvals, costs and programme impacts, every entry traced to a source document, with anything inferred rather than evidenced visibly markedThe account itself, and the conclusivity dates, which are the one thing a better record cannot save you from

Why does the record matter more than the analysis?

Because the analysis is used once and the record is used for years.

Referrals to adjudicator nominating bodies reached 2,264 in the year to April 2024, the highest level recorded and around 9 per cent up on the previous year (King's College London and the Adjudication Society). Every one of those is a retrospective reconstruction exercise: who knew, when, what information existed, what was instructed, what was decided and on what evidence.

A firm that has been assembling that record contemporaneously arrives at an adjudication with a file. A firm that has not arrives with two people and a weekend. The difference does not show up on any dashboard, and it is worth more than every hour the system saved.

What does this do to how junior surveyors learn?

It threatens the traditional route and offers a better one, and which of those a firm gets is a design decision rather than an outcome.

Traditionally juniors learn through repetition: measurement, applications, document review, reconciliation. Remove all of it and capability development has to be rebuilt deliberately. The model that works is not the junior receiving answers. It is the junior producing an assessment, the system challenging it against the evidence, the junior defending their reasoning, alternative interpretations being tested, and the decision being justified.

That is a harder day than either the old way or the lazy new way, and it develops commercial judgement faster than either. It also produces exactly the professional the RICS standard requires: one who understands the system's limitations, assesses reliability and keeps their own judgement at the centre. Related: tracking variations before they become disputes, JCT and NEC notices as conditions precedent and disputes are never born big.

Three disciplines, and in each of them the machine assembles and the professional decides. part 11, build an AI that watches work, never one that watches people asks what happens when you stop building assistants and start building oversight, and where the line is that HR, the workforce and the regulator will not let you cross.

What five questions should you be able to answer now?

Attempt these before tomorrow. Each has a defensible answer, and each is answered at the top of the next part.

  1. Why does a single general assistant watching a whole project fail, and what replaces it?
  2. What four things must every discipline agent have defined before it is built?
  3. What is the difference between observing artefacts and observing people, and why does the distinction decide whether a system can be sold?
  4. What is the most valuable table in the whole system, and why?
  5. What precision and recall targets would you set before letting an agent see live data?

Which sources is this part built on?

Every figure quoted above resolves to one of these. Each was checked before publication.

AI Metric is a construction-native AI consultancy. If your team is spending more time operating software than doing their job, book a 30 minute call.