Skip to content
AI Metric

Chris M.

Build an AI that watches work, never one that watches people

Oversight in a project team is not limited by knowledge. It is limited by attention, and attention is the one thing a machine has in surplus. What it is pointed at decides whether the system is an asset or a liability.

What were the answers to the previous five questions?

1. Why does a single general assistant watching a whole project fail, and what replaces it?

Because oversight is discipline specific. A three week steel delivery slip is a float question to the planner, a notice trigger to the commercial lead, a re-sequencing question to design and a standing time question on site. One agent asked to watch all of it produces generic observations nobody acts on, or so much noise it is switched off by week three. Narrow agents with defined remits replace it.

2. What four things must every discipline agent have defined before it is built?

A remit, a knowledge source, an escalation route, and a prohibition list. The prohibition list matters as much as the remit: the commercial agent never sets a valuation and never issues a notice, the HSE agent never authorises work or closes a safety action.

3. What is the difference between observing artefacts and observing people, and why does the distinction decide whether a system can be sold?

An artefact is a document, an email sent, a revision issued, an approval given, a value entered: business records created in the course of work. A click, a screenshot or time at keyboard is monitoring the person. The second is legally fraught and commercially fatal, and it will be stopped at the door by HR, legal or the workforce. Note what the distinction does not do: it is a test of proportionality, not a route out of data protection law, and a report that singles out an identifiable individual is personal data whatever the method is called.

4. What is the most valuable table in the whole system, and why?

The override table. Approvals confirm the system was used. Overrides tell you where it is wrong, where the knowledge is stale, or where the real process differs from the written one, and they are the raw material for deciding what to automate next.

5. What precision and recall targets would you set before letting an agent see live data?

Set them before the build and measure against a back test on a completed project. Catching 60 per cent of real issues at 90 per cent precision beats catching 95 per cent at 40 per cent precision, because false positives destroy adoption faster than false negatives destroy value.

Why is oversight the thing that gets displaced?

Because everyone on a project leadership team runs an oversight function on top of a delivery function, and delivery has a date on it.

The commercial lead is supposed to notice when a change was instructed verbally and never registered. The design manager is supposed to notice when a package is proceeding on information issued at a suitability code that does not permit construction. The planner is supposed to notice when an activity has quietly moved onto the critical path. They catch most of it. What they miss is almost never missed through incompetence. It is missed because the person who should have caught it was in a meeting about something else.

That is an attention gap rather than a knowledge gap, which matters because the two have different solutions. You cannot train your way out of an attention gap. You can add attention.

What does the architecture look like?

Three layers. Narrow discipline agents that observe and flag, a supervisory layer that deduplicates, challenges and routes, and a knowledge layer underneath: the controlled, current, project specific source material the agents reason from, which carries most of the implementation effort and almost none of the attention.

SUPERVISORY LAYEROrchestratorDeduplicates and ranksChallenge agentDifferent model lineageAuthority engineRules table, not a modelPattern engineOffline, finds repeatsDISCIPLINE AGENTS, EACH WITH A REMIT, A SOURCE, AN ESCALATION ROUTE AND A PROHIBITION LISTCommercialPlanningDesignProcurementQualityHSEProductionInformationKNOWLEDGE LAYER, WHERE MOST OF THE WORK ACTUALLY ISContract andamendmentsDelegatedauthority matrixCurrentrevisionsCorrespondencerecordCompany policyProject ontologyObserve artefacts, never people. If you could not describe the monitoring at an induction, the scope is wrong.
Eight discipline agents mapped onto roles that already exist, a supervisory layer above them, and the knowledge layer underneath that carries most of the implementation effort.
AgentWatchesNever
CommercialApplications, valuations, CVR movement, variation registers, notices, subcontract ordersSets a valuation or issues a notice
PlanningThe accepted programme against progress, logic, calendar and constraint changes between revisions, erosion of terminal float and time risk allowance as distinct from total float, milestone and Key Date exposure, and access or possession windowsReschedules, re-baselines, submits or accepts any programme, or determines the critical path for an entitlement
DesignRFI and TQ turnaround against the contractual period for reply, revisions and status codes, information release against procurement lead insReturns a review status on a design submission, states that a design complies, or uplifts a container status code
ProcurementPackage status against required on site dates, lead ins, scope boundaries between packagesPlaces an order
QualityITPs, hold points, test results, non conformance reports, handover evidenceCloses an NCR or signs off a hold point
HSEMethod statement currency against actual activity, permit expiry, inspection frequency, near miss clustersAuthorises work or closes a safety action
ProductionDiaries, allocation sheets, labour returns, delivery tickets, plant on hireWrites a record as fact without human confirmation
InformationContainer naming and metadata against the project BIM execution plan and the UK National Annex, distribution, revision integrity, archive containers still in circulationUplifts a status code, records a review as complete, or issues or supersedes a container

Above them sit four components, none of which is a discipline. An orchestrator takes an event, decides which agents it belongs to, deduplicates and ranks by consequence, so one steel delivery slip produces one ranked item rather than four competing ones. A challenge agent, which is the challenge function from parts 6 and 10 given a place in the architecture, reviews consequential drafts against contract, policy, prior correspondence and delegated authority, deliberately built on a different model family and prompt lineage from whatever produced the work. An authority engine maps a person to their delegated authority: a rules table, not a model, because that question is deterministic. A pattern engine runs offline against the ledger and never speaks to anyone in the moment.

What happens when it misses something?

It will, and every argument for the system has to survive that sentence. Sixty per cent recall at ninety per cent precision, which is the right target, means four real issues in ten are not flagged.

  • No manual check is removed on the strength of an agent. Not the inspection regime, not the permit checks, not the walk. It is additive to the arrangements in your safety management system, and standing a check down because the system now watches it is a change to those arrangements that goes through the same assessment as any other.
  • No flag is not a clearance. Put that sentence on the interface and in the induction, in those words. The failure mode is not that the system is wrong. It is that a supervisor stops looking because nothing appeared on a screen.
  • Measure the misses on a schedule. Precision falls out of the flag queue. Recall is only visible if somebody goes looking for what was not flagged, and nobody does that unless it is in the diary. Once a quarter, have a competent person review a completed period manually and count what should have been flagged and was not.

The same logic reaches the ledger. Every flag and every dismissal is disclosable, and a dismissed safety flag with no recorded reason is the worst single record this system can produce. Make the reason mandatory on anything from the health and safety or quality agents, and set the retention period deliberately rather than by default.

Where is the line that must not be crossed?

At the boundary between the work and the worker, and it is both a legal line and a commercial one.

ICO guidance on monitoring workers names keystroke monitoring specifically as processing likely to require a data protection impact assessment, alongside biometric data, monitoring that may result in financial loss, and profiling or special category data used to decide access to services. It says employers should seek and document the views of workers or their representatives before introducing monitoring, unless there is a good reason not to (ICO). Research the ICO commissioned alongside that guidance found around 70 per cent of the public surveyed said they would find workplace monitoring intrusive, which is a finding about expectations rather than about anyone experience of being monitored.

Be precise about what that distinction does and does not do, because firms get it wrong in a way that costs them. It is a sound test of whether monitoring is proportionate. It is not a route out of data protection law. A site diary has one author, and a report listing record gaps on days with delay events identifies the person who keeps that diary as surely as naming them would. Information that singles out an identifiable individual is personal data whatever the method is called, so there is still a lawful basis to establish, still a duty to tell people, and still an impact assessment to do before the build rather than after the first objection.

There is a commercial argument alongside the legal one, and it is the more persuasive of the two. Artefact level observation produces better flags. A click tells you somebody spent nine minutes in a spreadsheet. A revision history tells you a forecast changed by forty thousand pounds with no supporting evidence. Only one of those is worth a manager's attention, and only one of them will survive a conversation with a trade union representative.

What turns flagging into something worth paying for twice?

The ledger, and what you do with it.

Flagging has a ceiling. If the commercial agent flags the same missing evidence on the same monthly cycle for a year, it has done its job twelve times and improved nothing. The compounding value is that every flag, every decision on that flag, every override with a stated reason and every rework event becomes a row. Run a pattern engine over that weekly and you get an evidenced list of where the process is failing, ranked by frequency and consequence.

Three sources feed that candidate list, not one. The override rows, the pattern engine, and the workarounds recorded during discovery in part 4. The third is the same finding as the first two, arrived at before the system existed and at no cost, which is why part 4 asks for them.

  • The same three fields are re-keyed between the cost system and the forecast forty times a month by four people. Candidate: integration.
  • Subcontractor applications are rejected for the same missing evidence 60 per cent of the time. Candidate: a pre-submission check pushed out to the supply chain, not a better internal review.
  • Progress claimed on the programme has no corresponding site record on roughly one day in five. Candidate: link record capture to the progress update rather than train people harder.

Each candidate carries a frequency, a time estimate, a consequence band and a proposed control type: automate it, assist it, or challenge it. A human process owner approves the candidate, and only then does anything get built. Year one you buy oversight. Year two you buy the automation the oversight found.

Start with one discipline, read only, flagging to one reviewer, back tested against a completed project so precision is known before anyone relies on it. Do the impact assessment first and talk to the people who will be observed before the system observes anything. See AI agents for construction directors and when agents replace products.

All of which raises the question the finance director will ask, and should. part 12, stop measuring AI by time saved answers it without reaching for hours saved.

What five questions should you be able to answer now?

Attempt these before tomorrow. Each has a defensible answer, and each is answered at the top of the next part.

  1. Why is time saved a poor primary measure of AI value in construction?
  2. A report that took four hours now takes forty minutes and changes no decision. What has been achieved?
  3. What are the six categories worth measuring instead, and what does each one ask?
  4. Why must a baseline be captured before anyone knows a tool is coming?
  5. What are the three honest outcomes of a six week measurement cycle?

Which sources is this part built on?

Every figure quoted above resolves to one of these. Each was checked before publication.

AI Metric is a construction-native AI consultancy. If your team is spending more time operating software than doing their job, book a 30 minute call.