Same score, eight times the price
Construction prices work by the rate. This week four leading AI models scored within two points of each other and their rates ranged over eight times. One doubles its rate part-way through a long document. Two charge more and give you nothing for it.

- 2 points
- between the best of four leading models and the fourth, on the same test
- 8 times
- between the cheapest and the dearest of those same four
- 200k
- the document size at which the cheapest one doubles its rate
Shipped this week
MetaMuse Glimmer 30B10 Aug
SpaceXAIGrok Bot, Grok 4.611 Aug
SpotifyAI Persona badge11 Aug
OpenAIGPT-5.6 Sol pricing12 Aug
AnthropicText watermarking14 Aug
Z.aiGLM-5.314 Aug
The models drew level. The prices did not.
What happened. On 12 August, Artificial Analysis scored Grok 4.6 and GPT-5.6 Sol at 61, Claude Opus 5 at 63 and Fable 5 at 62. Anthropic also confirmed Claude Sonnet 5 will not rise in September. It stays at $2 and $10.
| Model | Score | Input per 1M | Output per 1M | On a long document |
|---|---|---|---|---|
| Claude Opus 5 | 63 | $5 | $25 | Same price throughout |
| Claude Fable 5 | 62 | $10 | $50 | Same price throughout |
| GPT-5.6 Sol | 61 | $5 | $30 | Higher rate of $10 and $45 |
| Grok 4.6 | 61 | $2 | $6 | $4 and $12 from 200k, on the whole request |
| Claude Sonnet 5 | Not tested here | $2 | $10 | Same price throughout |

Watch the long ones. Grok 4.6 is cheapest until your document passes 200,000 tokens. Then its rate doubles, on the whole request, not just the extra. Claude holds one price throughout.
The gap adds up. A thousand documents costs about $260 on Grok 4.6, $700 on Opus 5 and $1,400 on Fable 5, for the same work. Most firms never price it. They pick a default and live with it.
And the rate is not the bill. A cheap model that needs three goes at something can cost more than a dear one that gets it right first time.
Do this. Pick one job you would hand over. Run it on two models. Write down what you spent and whether you had to redo it. That is your real cost. Our note on choosing a model still stands.
Agents got a computer of their own
What happened. SpaceXAI launched Grok Bot on 11 August. Each agent gets its own cloud computer, signs into your apps, and keeps working when your laptop is shut. Show it a job once and it saves the steps. The day before, Meta released Muse Glimmer, small enough to run on one graphics card.


Why it matters. Automating a job used to mean somebody writing code for it. Now you show the agent. That puts it in reach of the person who does the work. It also means an agent is signed into your systems, which is a permissions question first.
| Tier | Example and rate | What it suits | Who checks it |
|---|---|---|---|
| Your own hardware | Muse Glimmer 30B, free to run | Sorting documents, reading a delivery ticket, filing | Spot checks |
| Cheap | GPT-5.6 Luna at $0.20 and $1.20 | Sorting an inbox, spotting which emails mention a delay | Spot checks |
| Everyday | Claude Sonnet 5 at $2 and $10, Grok 4.6 at $2 and $6 | A weekly report from site records, a first pass on an RFI log | A named person |
| Top | Claude Opus 5 at $5 and $25, GPT-5.6 Sol at $5 and $30 | A subcontract read against its amendments | A named person, every time |
| A human | Not a model | Anything priced, contractual or about safety | The duty holder |
Do this. List the five jobs that eat your week. Put each on a row of that table and write a name beside it. No name means not ready.
You can now tell when a model helped write something
What happened. Anthropic started marking Claude's text on 14 August. The mark is invisible, survives copying and light editing, and disappears if you rewrite the words. On 11 August Spotify announced a badge for AI-generated artists.

Why it matters. Your records already show who decided what, on what information. Now people will ask which tool helped. The answer is not detection software. It is writing it down at the time, as our Gateway 3 note describes.
Do this. Add two columns to your document register: which tool made the draft, and who checked it before it went out.
The common thread
None of this is a cleverer model. It is a price list with a trap in it, an agent you train by showing, and a mark that says a model helped but not who answers for it. The hard part has moved from what AI can do to what you ask it to do, who checks it, and whose name goes on it. That is a scope of works, and construction knows how to write one.

To run any of this on your own numbers, say so.
Sources
Every figure and price in this edition was read from the page linked below on 15 August 2026. Where a number could not be confirmed at its publisher, it is not in the edition.
- Artificial AnalysisGrok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency
- SpaceXAIModels and pricing
- SpaceXAIIntroducing Grok Bot
- AnthropicClaude Platform pricing
- AnthropicHow Claude's text watermarking works
- OpenAIAPI pricing
- Meta AI ResearchIntroducing Muse Glimmer: an open agentic model that runs on your device
- SpotifyIntroducing a new label for AI-generated artist identities on Spotify
- SiliconANGLEZ.ai debuts GLM-5.3 with long-horizon coding, cybersecurity upgrades
Followed this week
Where the week was picked up, and where the screenshots on this page were captured from. Every figure and price above was then checked at the publisher and drawn from that source, so no chart here is traced off a video.
- Matthew BermanxAI actually did it... (Grok 4.6)
- Matthew BermanAI News: ChatGPT Ultrafast, Grok 4.6, 3 New Open-Source Models, and more!
- Matthew BermanMark Zuckerberg just called out Dario (and Anthropic)