Skip to content
AI Metric

Chris M.

Speed just went on the rate card

A mobile crane on hire costs the same whether it is lifting or standing. A monthly cost report does not care if it waits until Monday. Every job runs both kinds of work at once. Until this summer AI was sold at one price and one speed, so nothing on the invoice ever told the two apart. OpenAI has now split it into three things you choose between: where the work runs, how fast the answer comes back, and how much thinking you are paying for.

14x
how much faster the top model now runs, for a few customers
2x
what a faster answer adds to the bill
39x
between the cheapest way to read a thousand documents and the dearest

The work stops living on one person's laptop

What happened. OpenAI is buying a company called Ona, announced on 11 June. Ona runs AI tools inside a firm's own secure cloud, so a job carries on after somebody shuts their laptop. OpenAI says the work worth doing now runs for hours or days rather than minutes. More than five million people use its Codex tool every week.

Why it matters. Take the lifting file on any decent sized job: the lift plans, the thorough examination certificates for every chain, sling and shackle, the RAMS, the appointed person's sign offs. Plenty of people have already pointed AI at that pile to sort it, check what is missing and chase it. It works. Then it turns out the whole thing runs on one person's laptop. It stops the week they are on leave. The record it built walks out of the door when they change firms.

The list OpenAI is selling this on is the list you would put in a subcontract: where it runs, what it is allowed to see, who holds the logins, what gets recorded, and who signs the output off. Those questions are not new. It is only new that you can now ask them of an AI tool and get a straight answer.

Do this. Write down every AI job your team leans on. Next to each one, put the machine it runs on and the name of somebody else who can start it again. One name in both columns is what to fix before you spend a penny on speed.

You can now pay extra for a faster answer

What happened. On 13 August OpenAI showed Ultrafast, its top model running up to 14 times faster than normal, which is quicker than anyone can read it. That one is with a handful of customers and not on sale. What you can buy today is Fast mode. Through the API it costs twice the normal rate. Inside Codex it burns credits two and a half times as fast.

What you are buyingHow fastWhat it costsWhen it is worth it
BatchSlower, queuedHalf the normal rateNobody is waiting
StandardNormalThe rate cardMost work
Fast mode, APIUp to 2.5 timesTwice the rateSomeone is waiting on the phone
Fast mode, Codex1.5 timesCredits go 2.5 times fasterSomeone is sat watching it
UltrafastUp to 14 timesNot publishedNot on sale yet
OpenAICodex speed settings, read 25 August
OpenAI's Codex documentation page titled Speed, stating that Fast mode increases supported model speed by 1.5x, that GPT-5.6 and GPT-5.5 consume credits at 2.5x the Standard rate and GPT-5.4 at 2x, and that API Priority processing costs 2x the Standard API token rate for GPT-5.6.
Read the second paragraph twice. The speed goes up by half. The meter runs at two and a half times. So hurrying a job in Codex costs about 1.7 times as much for the same result.

Why it matters. Here is where that choice gets made. The crane is on site, the slinger is booked, and the unit that turns up is not the one on the lift plan. The appointed person needs to know what the plan and the RAMS said about that load at that radius. Until they do, the crane, the gang and the delivery are all stood there. That is charged by the hour and it is the sort of hour a project manager remembers for a year.

That answer is worth paying double for. The same tool reading through the lifting file overnight to tell you which certificates expire next month is not. That one can queue, and it runs at half price while everybody is asleep.

Do this. Name the two jobs in your week where somebody is genuinely stood waiting on an answer. Pay for speed on those two. Put the rest on the overnight rate and stop paying a premium for work nobody is watching.

The cheap model got 80 per cent cheaper

What happened. On 30 July OpenAI cut the price of Luna, its cheapest model, by 80 per cent, and Terra, its middle one, by 20 per cent. The day before, it explained where the money came from. It used its own top model to rewrite the software that runs the models. That took 20 per cent off the cost of running them, and made them produce text more than 15 per cent more efficiently.

$0$200$400$600$800$1000$1200SPEND ON 1,000 DOCUMENTS, USD$28.80$57.60GPT-5.6 Luna$288$576GPT-5.6 Terra$560$1,120GPT-5.6 SolStandardFast modeHOW TO READ THISEach pair is the same thousand documents. Solid is the normal rate, dashed is Fast mode.Cheapest model at normal speed against the dearest in a hurry: about 39 times.MODELLED, NOT MEASURED. 120,000 input and 4,000 output tokens per document,at the published short context rates, read 25 August 2026.
Two decisions, not one. Which model reads the document, and whether anybody is stood waiting while it does.Source: OpenAI API pricing, standard and Fast mode, short context rates, read 25 August 2026
OpenAIAPI pricing, flagship models, read 25 August
OpenAI's API pricing page, flagship models table, showing short context prices per million tokens: gpt-5.6-sol at $4.00 input and $20.00 output, gpt-5.6-terra at $2.00 and $12.00, and gpt-5.6-luna at $0.20 and $1.20.
Three versions of the same model, priced before anything is added for speed. Sol costs twenty times what Luna costs. It is worth that on the jobs where being wrong is expensive, and it is not worth it on filing.

Why it matters. Count what is in one lifting file, or in a subcontractor's O&M pack, and a thousand documents stops sounding like a lot. Reading that pile costs about $29 on the cheapest model and about $1,120 on the dearest one in a hurry. Same documents, same work, same answer at the end of it.

Which one you need depends on the job, not the budget. Pulling the expiry date off a certificate and filing it under the right crane is well defined work, and the cheap model does it. Reading a lift plan against a changed load, or a subcontract against its amendments, is not well defined, and that is what the expensive one is for. Most business cases were priced on one number, before July, and nobody has been back to them since.

Do this. Take one job you turned down on cost, price it again at today's rates, and check which model it actually needs. If the answer is only ever checked by a person anyway, it does not need the dear one.

What to ask before you sign

An AI quote now has three questions in it: where the work runs, how fast the answer comes back, and how much thinking you are paying for. Those are the same three things you would pin down on any package before you let it start, and they belong in the paperwork for the same reason.

If somebody is asking you to approve an AI tool this quarter and the paper in front of you does not answer all three, that is a twenty minute conversation.

Sources

Every figure and price in this edition was read from the page linked below on 25 August 2026. Where a number could not be confirmed at its publisher, it is not in the edition.

Followed this week

Where the week was picked up. Every figure and price above was then checked at the publisher and drawn from that source, so no chart here is traced off a video.

This series reads the week’s AI releases from a construction delivery position, not a technology one. If one of these lands on a package you are running and you want to talk about what it would take, book a 30 minute call.

Commentary on publicly announced capability, written from twenty-plus years of Tier 1 delivery. It is not design guidance, fire or building safety advice, or contractual advice, and no live project is described.