Skip to content
AI Metric

Chris M.

It stopped going past what it was asked

The question that stops an AI tool getting near a live job is never whether it is clever. It is what happens on the run where it misreads the brief and carries on anyway, at two in the morning, in a system somebody has to answer for. This week that question got a published number against it.

Written up on 28 September, covering the week of 4 September. Every figure below was read at the publisher on 28 September.

The scope number is the one to read first

What shipped. OpenAI released GPT-6 Astra on 3 September, generally available the following day. Alongside the capability claims they published a scope test: given a task that is difficult or impossible, does the model go beyond the target it was authorised to touch. Their previous model did, in 48 per cent of cases without production safeguards. Astra did in none.

Where it lands. That is the number a contractor should care about, and it is buried under the benchmark scores. An agent allowed to read a project folder and update a register is useful. The same agent deciding on its own that the neighbouring folder looks relevant is an incident, and on a live project it is an incident somebody has to disclose. Nought per cent is not a guarantee on your systems with your data. It is the first time the failure has been measured and published at all.

Monday morning. If anyone in your business is piloting an agent, write down today what it may read and what it may change, before you look at any model. The scope is yours to set and no release will set it for you.

WENT PAST WHAT IT WAS ASKEDshare of attempts, without production safeguards48%GPT-5.6 Sol0%GPT-6 AstraOSWORLD 2.0, LATENCY SIMULATIONminutes per task, and the score it got there withGPT-5.6 Solabout 75 min65.7%GPT-6 Astraabout 40 min72.6%HOW TO READ THISLeft: how often each model went beyond the target it was given, in a test built for that question.Right: the same workload, in minutes, with the score beside it so speed is not read on its own.Both read from OpenAI's own GPT-6 Astra announcement, 4 September 2026. Their tests, their figures.
Two measures that only mean anything together. A tool that finishes in half the time and exceeds its brief in half the runs has saved nobody anything.Source: OpenAI, GPT-6 Astra announcement, read 28 September 2026

Faster and better at once, which is new

What shipped. On OSWorld 2.0, a test of long computer-use workflows, Astra scored 72.6 per cent at roughly 40 minutes per task against its predecessor's 65.7 per cent at roughly 75 minutes. About 47 per cent less time for a higher score.

Where it lands. Until now these moved against each other: the careful setting was slow and the quick setting was worse. Work that was not worth automating because it took the machine longer than the person is worth pricing again. Collating a handover pack, chasing missing certificates through a folder tree, reconciling a delivery log against a register: all tedious, all checkable, all now inside a working morning rather than a working week.

Monday morning. Pick the job in your week you would never give to a graduate because explaining it takes longer than doing it. That is the shape of task this changed, and it is the one worth timing honestly.

OpenAIthe launch, from the publisher
The vendor making its own case, which is worth watching for what it chooses to lead with. The scope result is not the headline here; it is in the written announcement. Introducing GPT-6 Astra: the most intelligent and aligned model in the world., published by OpenAI on YouTube.

It can now search its own toolbox

What shipped. The model card lists what Astra can call: web search, file search, computer use, a hosted shell, the Model Context Protocol, and a tool search that lets it pick from its own available tools rather than being handed a fixed list. The window is 1,050,000 tokens, with a knowledge cutoff of 30 April 2026.

Where it lands. Searching stopped being something you do before you ask and became a step inside the model's own loop. That is why the cutoff matters: the gap between what a model knows and what is true on your job is exactly the gap a search is there to close, and it can only close it against information it is allowed to reach.

Monday morning. Ask where your project information actually lives. A model that can search everything is worth nothing if the record of the job is on four people's phones.

OpenAIGPT-6 Astra model documentation
The GPT-6 Astra model page on OpenAI's developer documentation, showing a price of ten dollars input and fifty dollars output per million tokens, a 1,050,000 token context window, 128,000 maximum output tokens, a knowledge cutoff of 30 April 2026, and a pricing row giving cached input at one dollar and cache writes at twelve dollars fifty.
The model card, which is duller and more useful than the announcement. The context window, the cutoff and the cached input rate are the three numbers that decide what this costs to run in anger.
What it can callWhat that meansWhere it bites on a project
Web searchLooks things up itself, mid-taskAnswers can be current, and can also be confidently wrong
File searchReads a body of documents you point it atOnly as good as what is actually filed
Computer useDrives applications and a browserNeeds a written boundary before it gets one
Model Context ProtocolConnects to your own systems through a standardScope it to a folder, never to the estate

What to write down before you let one near anything

The releases are ahead of most people's governance, and the gap is not technical. It is that nobody has written down what the thing may touch.

If you are being asked to approve an agent this quarter and the paper in front of you does not say what it may read, what it may change and who signs off, that is a twenty minute conversation.

Matthew Bermanthe week, followed
Where this week got noticed. The read is his; the figures above were taken off OpenAI's own pages rather than from any commentary, which is the rule here. ASTRA IS HERE (GPT-6 RELEASED), published by Matthew Berman on YouTube.

Sources

Every figure and price in this edition was read from the page linked below on 28 September 2026. Where a number could not be confirmed at its publisher, it is not in the edition.

Followed this week

Where the week was picked up. Every figure and price above was then checked at the publisher and drawn from that source, so no chart here is traced off a video.

This series reads the week’s AI releases from a construction delivery position, not a technology one. If one of these lands on a package you are running and you want to talk about what it would take, book a 30 minute call.

Commentary on publicly announced capability, written from twenty-plus years of Tier 1 delivery. It is not design guidance, fire or building safety advice, or contractual advice, and no live project is described.