Classify by consequence, not by which logo is on the screen
The instinct is to treat consumer tools as dangerous and enterprise tools as safe. That sorts by data security and ignores correctness entirely, which is the half that produces the expensive failures.
What were the answers to the previous five questions?
1. Two people use the same AI product. One summarises internal minutes, the other drafts contractual notices. Should they carry the same controls, and why?
No. The product is identical and the exposure is not. A poor internal summary costs a clarifying email. A defective notice can lose an entitlement or start time running. Controls follow consequence, so the two sit in different bands despite sharing a licence.
2. What two variables should determine whether a use case is Green, Amber or Red?
The consequence if the output is wrong, and how late that error would surface. High consequence caught immediately in normal review is manageable. Moderate consequence discovered on site or in an adjudication is not.
3. Why is "is the tool enterprise grade" the wrong basis for classification?
Because it answers where the data sits, not whether the answer is right. An enterprise secure system will still work from a superseded revision, miss a project specific amendment or state responsibility it cannot evidence. Security is a precondition, not a classification.
4. Which kind of use case feels low risk but should be classified Red?
Summarising a design team meeting where a fire strategy assumption changed. It looks like note taking. The summary becomes the record of a decision affecting statutory compliance, and the error would not surface until building control or later. Red, on both variables at once.
5. What is the minimum control that should attach to any Red classified use?
Four are non-negotiable: a named decision owner, explicit approval by an authorised role working from the underlying evidence rather than the AI summary, a defined escalation route, and a retained record of what was reviewed and decided. The band table adds three more that apply once the use case is running: defined delegated authority, formal testing before deployment, and periodic assurance afterwards.
Why does the two variable model work better than a list of tools?
Because a list of tools goes out of date at the next software update, and because it sorts on the wrong attribute. Consequence and lateness are stable properties of the work itself.
Lateness deserves more weight than it usually gets. Construction has spent decades building gates precisely because the cost of an error rises steeply with the distance between the mistake and its discovery. An AI error caught by the person who prompted it costs nothing. The same error discovered when the cladding is up costs whatever it costs.
What sits in each band?
Green is assistance where being wrong costs a correction. Amber is anything that materially informs professional work. Red is anything where being wrong reaches safety, statute, certification or a material sum, and the control steps up at each boundary.
| Band | Typical use | Minimum control |
|---|---|---|
| Green | Grammar, formatting, agendas, brainstorming, general research, low risk internal summarising | Normal professional review. Approved system, appropriate data. |
| Amber | Contract analysis, programme analysis, tender review, specification review, technical interrogation, client reporting, CVR challenge, project correspondence | Competent human review, source verification, assumptions visible, named decision owner, audit trail where proportionate. |
| Red | Safety critical analysis, statutory compliance, design acceptance, formal certification, significant payment decisions, material contractual positions, any notice operating as a condition precedent, any agent able to act on external systems | Named decision owner, explicit approval against underlying evidence, defined delegated authority, strong audit trail, formal testing, escalation route, periodic assurance. |
Note what is absent from all three: any mention of a product. The same assistant can legitimately appear as Green for drafting an agenda, Amber for interrogating a specification and Red for anything touching a gateway submission under the Building Safety Act. One register, three rows, three sets of controls.
Where do firms get the classification wrong?
In four recurring places, all of them in the same direction.
- Meeting minutes. Filed as Green because it is note taking. Minutes are frequently the only record that a decision was made, by whom, and on what basis. If a summary can become the evidence of a commercial decision, it is Amber. If it can become the evidence of a decision bearing on statutory compliance or safety, a design team meeting where a fire strategy assumption moved being the standard case, it is Red. One activity, one tool, three bands, and the two variables decide which.
- Email refinement. Filed as Green because it is writing help. Outbound project correspondence is a project record with contractual weight, as part two set out. Internal chat is Green, correspondence to the other side is Amber.
- Anything involving a graduate. Firms classify by the seniority of the user rather than by the consequence of the output. An assistant quantity surveyor producing a commercial assessment is not exercising authority at all, because the authority sits with whoever signs. What has changed is that the signatory is now approving work that reads as though it were senior, so the review requirement rises and the reviewer has to work from the evidence rather than from the document in front of them.
- Embedded features. AI that arrived inside estimating or design software by update often gets no classification at all because nobody procured it. It is doing exactly the work you would classify Amber if you had noticed.
How do you make the classification survive contact with a busy week?
By making the control proportionate enough that nobody has an incentive to route around it. A three day approval queue for a meeting summary teaches people to stop using the approved route, and you are back to a register you cannot believe.
| Gate | Applies to | Mechanism |
|---|---|---|
| Self check | Amber that stays inside the business: internal reports, drafts | User reviews their own output against a short checklist. No second person. |
| Competent review | Amber that leaves the business: contract interpretation, tender responses, client reporting | A named competent person reviews before it is used or sent. |
| Named approval | Red: safety, statutory compliance, formal certification, major payment | Authorised role reviews the underlying evidence, not the summary, and formally approves. |
Write the gate into the register row itself rather than into a separate policy document. A control that lives next to the use case it governs gets followed. A control that lives in a policy on SharePoint gets acknowledged once at induction. Related: controlled AI adoption not blanket bans and private AI versus public chatbots.
Classification tells you how much human control a use case needs. It does not tell you what that control consists of, and "a human checks it" is not an answer. part 6, what meaningful human involvement now has to mean makes it one.
What five questions should you be able to answer now?
Attempt these before tomorrow. Each has a defensible answer, and each is answered at the top of the next part.
- Since 5 February 2026, what does UK law say makes a decision "solely automated"?
- What four attributes does a human reviewer need before their involvement counts as meaningful?
- What is automation bias, and what organisational conditions make it worse?
- "A competent person reviews it" is not a control. What would a written control have to specify before a reviewer at five o clock on a Friday knows exactly what to look at?
- Why should a firm log overrides as carefully as it logs approvals?
Which sources is this part built on?
Every figure quoted above resolves to one of these. Each was checked before publication.