Skip to content
AI Metric

Chris M.

Same score, eight times the price

Construction prices work by the rate. This week four leading AI models scored within two points of each other and their rates ranged over eight times. One doubles its rate part-way through a long document. Two charge more and give you nothing for it.

Steel frame under erection, cellular beams and columns against a blue sky.
Every piece here was priced by the rate before anyone lifted it. AI is sold the same way, and this week the rate stopped being one number.
2 points
between the best of four leading models and the fourth, on the same test
8 times
between the cheapest and the dearest of those same four
200k
the document size at which the cheapest one doubles its rate

Shipped this week

  • MetaMuse Glimmer 30B10 Aug
  • SpaceXAIGrok Bot, Grok 4.611 Aug
  • SpotifyAI Persona badge11 Aug
  • OpenAIGPT-5.6 Sol pricing12 Aug
  • AnthropicText watermarking14 Aug
  • Z.aiGLM-5.314 Aug

The models drew level. The prices did not.

What happened. On 12 August, Artificial Analysis scored Grok 4.6 and GPT-5.6 Sol at 61, Claude Opus 5 at 63 and Fable 5 at 62. Anthropic also confirmed Claude Sonnet 5 will not rise in September. It stays at $2 and $10.

6061626364$0$10$20$30$40$50ARTIFICIAL ANALYSIS INTELLIGENCE INDEXAXIS STARTS AT 60. THE WHOLE FIELD IS TWO POINTS WIDE.OUTPUT PRICE, USD PER MILLION TOKENSTHE FRONTIERGrok 4.6$6 out, index 61Claude Opus 5$25 out, index 63GPT-5.6 Sol$30 out, index 61Claude Fable 5$50 out, index 62HOW TO READ THISLeft is cheaper. Higher scores better. The index is one combined score from independent testing.Best score: Claude Opus 5 on 63. Cheapest at the same score as Sol: Grok 4.6, 61 for $6.Dearest at that score: GPT-5.6 Sol, 61 for $30. Anything in the shaded area is beatenon price and score at once by the model at its corner.
Two of these four are beaten on score and price at the same time by another model on the chart.Source: Artificial Analysis, 12 August 2026; vendor pricing pages, read 15 August 2026
ModelScoreInput per 1MOutput per 1MOn a long document
Claude Opus 563$5$25Same price throughout
Claude Fable 562$10$50Same price throughout
GPT-5.6 Sol61$5$30Higher rate of $10 and $45
Grok 4.661$2$6$4 and $12 from 200k, on the whole request
Claude Sonnet 5Not tested here$2$10Same price throughout
SpaceXAIGrok 4.6 evaluation table, published 11 August
A benchmark table published by SpaceXAI comparing Grok 4.6 High, Grok 4.5 High, GPT-5.6 Sol Max and Fable 5 Max across the AA Intelligence Index, GDPVal-AA v2, CursorBench, DeepSWE, FrontierCode, APEX-Agents, Terminal-Bench, APEX-SWE, AA-Briefcase and Harvey LAB.
Captured from Matthew Berman on YouTube. Videos in the sources.
The vendor's own comparison. Read it as a claim rather than a finding: bold marks the best score in each row, and SpaceXAI's own footnote says the rival numbers are whatever those vendors have reported publicly. The AA Intelligence Index row is the exception, because that number is somebody else's, and it agrees with the chart above.

Watch the long ones. Grok 4.6 is cheapest until your document passes 200,000 tokens. Then its rate doubles, on the whole request, not just the extra. Claude holds one price throughout.

$0$1$2$3$4$50k100k200k300k400k500kEFFECTIVE INPUT RATE, USD PER MILLION TOKENSPROMPT SIZE, TOKENSClaude Opus 5, flat to 1M contextGrok 4.6doubles, on the whole request200kRoughly 200,000 tokens of a live jobOne subcontract with a bespoke amendment schedule, its appendices, and the RFI log it is being read against.HOW TO READ THISA token is a chunk of text, roughly three quarters of an English word.Each line is what that model charges per million tokens of what you send it,as the document you send it gets longer.
A subcontract read against its amendment schedule and RFI log is about the size that trips this.Source: SpaceXAI and Anthropic pricing documentation, read 15 August 2026

The gap adds up. A thousand documents costs about $260 on Grok 4.6, $700 on Opus 5 and $1,400 on Fable 5, for the same work. Most firms never price it. They pick a default and live with it.

$0$250$500$750$1000$1250$150002505007501000CUMULATIVE SPEND, USDDOCUMENTS READClaude Fable 5$1400Claude Opus 5$700Claude Sonnet 5$280Grok 4.6$264HOW TO READ THISFollow a line right to the number of documents you would really process in a year,then read the spend off the left. All four lines do the same work.MODELLED, NOT MEASURED. One read = 120,000 input and 4,000 output tokens,at each published rate.
Same job, four bills. The choice behind the gap is usually made once, early, by whoever set up the account.Source: Modelled from published rates. Assumption printed on the figure.

And the rate is not the bill. A cheap model that needs three goes at something can cost more than a dear one that gets it right first time.

TURNS TO RESOLVE A TASKINPUT TOKENS CONSUMEDGrok 4.6~53~0.5BClaude Opus 5~103~2.0BHOW TO READ THISA turn is one exchange between the agent and the model. A long job takes many.Half the turns and a quarter of the text, for the same result on the same test.A rate card cannot tell you this. Only a finished job can.
Grok 4.6 finished the same test in half the steps and a quarter of the text.Source: Artificial Analysis AA-Briefcase, 12 August 2026

Do this. Pick one job you would hand over. Run it on two models. Write down what you spent and whether you had to redo it. That is your real cost. Our note on choosing a model still stands.

Agents got a computer of their own

What happened. SpaceXAI launched Grok Bot on 11 August. Each agent gets its own cloud computer, signs into your apps, and keeps working when your laptop is shut. Show it a job once and it saves the steps. The day before, Meta released Muse Glimmer, small enough to run on one graphics card.

A site office at dusk, monitors lit, hard hat on the desk, nobody at the chair, a lit frame and tower crane through the windows.
The point of an agent with its own computer is the empty chair. The work carries on after everyone has gone.Image generated with AI. The mark in the corner is the generator's own.
SpaceXAIGrok Bot product page, 11 August
The Grok Bot product page, showing a messaging-style interface with named agents down the left and an agent signing into an application on its own computer.
Captured from Matthew Berman on YouTube. Videos in the sources.
It looks like a messaging app because that is the point. The agents down the left each have a job, and the panel on the right is one of them signing into an application on a computer of its own.

Why it matters. Automating a job used to mean somebody writing code for it. Now you show the agent. That puts it in reach of the person who does the work. It also means an agent is signed into your systems, which is a permissions question first.

PROJECT TEAMsets the objectiveORCHESTRATORdecides who does what, and on which modelPROGRAMMERead: programme,progressNo client commsCOMMERCIALRead: valuations, CVRNo send, no submitDOCUMENT CONTROLRead and file:drawingsNo issue to clientSITE RECORDSRead: diaries, photosNo noticesCDE, EMAIL, FINANCE SYSTEM, SITE APPS, SIGNED-IN BROWSERNAMED HUMAN APPROVESbefore anything reaches a client, a subcontractor or the record
Vendors draw the top three rows. The two that decide whether this is safe on a live job are the ones nobody draws.
TierExample and rateWhat it suitsWho checks it
Your own hardwareMuse Glimmer 30B, free to runSorting documents, reading a delivery ticket, filingSpot checks
CheapGPT-5.6 Luna at $0.20 and $1.20Sorting an inbox, spotting which emails mention a delaySpot checks
EverydayClaude Sonnet 5 at $2 and $10, Grok 4.6 at $2 and $6A weekly report from site records, a first pass on an RFI logA named person
TopClaude Opus 5 at $5 and $25, GPT-5.6 Sol at $5 and $30A subcontract read against its amendmentsA named person, every time
A humanNot a modelAnything priced, contractual or about safetyThe duty holder

Do this. List the five jobs that eat your week. Put each on a row of that table and write a name beside it. No name means not ready.

You can now tell when a model helped write something

What happened. Anthropic started marking Claude's text on 14 August. The mark is invisible, survives copying and light editing, and disappears if you rewrite the words. On 11 August Spotify announced a badge for AI-generated artists.

WHAT HAPPENS TO THE TEXTWHAT ANTHROPIC SAYS THE MARK DOESCopied and pastedMark travels with the textLightly editedMay persist through some editingShort extractWeaker: detection falls on small samplesRewritten word for wordGoneSo a mark found in a document is evidence that a model processed the text.A document with no mark is not evidence that none did. Absence proves nothing.
A mark that is found tells you a model touched the text. A mark that is missing tells you nothing at all.Source: Anthropic, How Claude's text watermarking works, 14 August 2026
AnthropicHow Claude marks AI-generated content, 14 August
Anthropic's help page titled How Claude marks AI-generated content, setting out marking commitments under the EU AI Act's Article 50 Code of Practice.
Captured from Matthew Berman on YouTube. Videos in the sources.
Worth reading in the original rather than in a summary of it. The commitments are specific, and so are the limits: marking applies from a stated date, across stated products, and Anthropic says plainly where it does not hold.

Why it matters. Your records already show who decided what, on what information. Now people will ask which tool helped. The answer is not detection software. It is writing it down at the time, as our Gateway 3 note describes.

Do this. Add two columns to your document register: which tool made the draft, and who checked it before it went out.

The common thread

None of this is a cleverer model. It is a price list with a trap in it, an agent you train by showing, and a mark that says a model helped but not who answers for it. The hard part has moved from what AI can do to what you ask it to do, who checks it, and whose name goes on it. That is a scope of works, and construction knows how to write one.

A city waterfront at dusk from the air, towers, docks and a river under a red sky.
Everything here leaves a record that has to answer for itself years later. That is the part no model can take on for you.

To run any of this on your own numbers, say so.

Sources

Every figure and price in this edition was read from the page linked below on 15 August 2026. Where a number could not be confirmed at its publisher, it is not in the edition.

Followed this week

Where the week was picked up, and where the screenshots on this page were captured from. Every figure and price above was then checked at the publisher and drawn from that source, so no chart here is traced off a video.

This series reads the week’s AI releases from a construction delivery position, not a technology one. If one of these lands on a package you are running and you want to talk about what it would take, book a 30 minute call.

Commentary on publicly announced capability, written from twenty-plus years of Tier 1 delivery. It is not design guidance, fire or building safety advice, or contractual advice, and no live project is described.