Construction finally has a number to compare itself against
Ask a project team whether their programme is performing and you get an opinion. There has never been an outside number to hold it against, so "we are doing about as well as anyone" has been impossible to challenge and impossible to prove. Three releases this month change that, and none of them is a capability announcement.
Buildots published the benchmarks instead of selling them
What shipped. Buildots opened its Intelligence Lab on 25 June, publishing benchmarks built from anonymised data across its client projects. Schedule adherence by sector, output rates by trade, and where in an activity the time actually goes. It is free to read, with no licence and no sales call in front of it.
Where it lands. The numbers are uncomfortable in a useful way. Buildots puts schedule adherence at around 65% on healthcare, roughly 57% on data centres, the low-to-mid 40s on commercial and industrial work, and under 39% on education. It also reports that the final 20% of an activity eats 27% of its duration. Every PM knows the last stretch of a package drags. Nobody could previously put a figure on it in a progress meeting and expect it to survive challenge. That changes what a delay narrative can be built on, because a structural pattern in the data is a different argument from an assertion about one subcontractor.
Monday morning. Take your last completed package and work out what proportion of activities finished in the week they were planned to. Compare it with the sector figure. If you are at 45% and the benchmark is 57%, you have a conversation with evidence in it rather than a feeling.
The price of reading every drawing revision fell again
What shipped. Google released Gemini 3.6 Flash on 21 July at $1.50 per million input tokens and $7.50 per million output, with a smaller Flash-Lite tier at $0.30 and $2.50. Google says the newer model needs about 17% fewer output tokens than its predecessor to do the same work, which is its claim rather than an independent finding.
Where it lands. Cost, not capability, is what has kept AI off the boring documents. A model good enough to read a 200-page subcontract and flag which clauses are conditions precedent has existed for a while. Running it across every revision of every drawing, every RFI and every day's site diary on a 40,000 sq m warehouse envelope package is an arithmetic question, and the arithmetic has been the blocker. At these rates, reading the whole record rather than the parts someone remembered to check moves from a business case into a line item. Note the free tier is not the same product commercially: on it, Google uses the content to improve its own models, which rules it out for anything priced or contractually sensitive.
Agents that keep working after you leave the office
What shipped. Anthropic extended Claude Cowork to web and mobile in beta on 7 July. Tasks now run in the cloud and continue with the laptop closed, and can be scheduled to run unattended. It is included on Pro, Max, Team and Enterprise plans, and is aimed at ordinary knowledge work rather than software engineering.
Where it lands. The constraint on this kind of automation in construction has never been the model. It is that the people with the worst admin load are the ones least likely to be sitting at a desk when the work needs doing. A site manager writing up at seven in the evening is not going to keep a laptop open while something processes. Assembling a weekly report from a dozen sources overnight and finding it done is a different proposition from starting the same task and babysitting it. Worth being clear-eyed about the beta label: unattended runs are exactly where a wrong output does damage quietly, so this belongs on internal reporting before it goes anywhere a client sees.
What the three have in common
None of these is a smarter model. They are a free measuring stick, a lower price per document, and work that continues when you put your phone in your pocket, which together say the limit on using this stuff in construction has stopped being what the technology can do. If you want to test that on your own numbers, the schedule-adherence exercise above takes an afternoon and needs no software at all, and we are happy to look at the result with you at ai-metric.com/contact/.