Chris M.Updated
The cheap model won on the work that pays
Most AI business cases in construction were priced on one assumption: the good model is dear, so you use it sparingly and put up with the cheap one everywhere else. That assumption stopped holding this week, and it stopped holding on the vendor's own figures rather than ours.
Written up on 28 September, covering the week of 22 September and GPT-6 Astra, released on 3 September. Every figure below was read at the publisher on 27 or 28 September.
Half the price, and it came top
What shipped. OpenAI released GPT-6 Sol and GPT-6 Luna on 22 September, cutting the price of both by half against what the previous generation charged. Sol went from four dollars to two per million input tokens, Luna from twenty cents to ten. On AutomationBench, a test of business workflows across 47 tools, Sol scored 33.2 per cent at 27 cents a task. Claude Opus 5 scored 26.9 per cent at 11.1 times that cost.
Where it lands. The cheap tier is now the right home for most of what a contractor would actually automate: reading a delivery ticket, pulling dates off certificates, turning a week of records into a draft narrative. That work is high volume and checkable, which is exactly the shape the cheap tier handles. Anyone still paying flagship rates to summarise documents is paying for reasoning they are not using.
Monday morning. Take one process you priced on last quarter's rates and price it again on this week's. The cheapest tier is a tenth of what it cost in July.

Anthropic cut its own price the same day
What shipped. Claude Opus 5.5 landed on 22 September at four dollars input and twenty dollars output, about 40 per cent cheaper on typical workloads than Opus 5 and generating output more than 30 per cent faster. Reading a cached token fell from fifty cents per million to twenty. Anthropic report a tester completing a 680,000 line code migration in under a day.
Where it lands. Two vendors cutting price on the same day is a market telling you something: the margin has moved from the model to what you do with it. For a business buying AI, that means a contract signed on today's rates is a worse deal in ninety days, and the tools should be swappable rather than welded in.
Monday morning. Check whether anything you are being sold this quarter locks you to one model. If it does, ask what happens at the next price cut, because there will be one.
| Model | Input | Output | Where it earns its place |
|---|---|---|---|
| GPT-6 Astra | $10 | $50 | The judgement call at the end, not the volume |
| Claude Opus 5.5 | $4 | $20 | Long agentic work, computer use, research |
| GPT-6 Sol | $2 | $10 | Multistep work that gets checked anyway |
| GPT-6 Luna | $0.10 | $0.50 | Reading, extracting, filing, at volume |
Caching became the biggest lever on the bill
What shipped. Less visible than the price cuts and worth more. Cached input on GPT-6 is discounted by 90 per cent, and OpenAI report that GitHub cut the share of prompt tokens needing fresh processing by more than half across billions of requests. Opus 5.5 reads cached tokens at a fifth of what Opus 5 charged.
Where it lands. An agent does not answer once. It plans, looks things up, calls a tool, checks itself and goes round again, so the bill is the number of passes multiplied by what each pass burns. Most of that is the same unchanging context every time: the contract, the register, the standing instructions. Structured so it is reused, it is nearly free. Structured badly, you pay full price for it on every lap.
Monday morning. If you are running anything in production, ask whoever built it what your cache hit rate is. If nobody knows, that is the cheapest performance work available to you this month.
Search moved inside the model’s own loop
Searching stopped being the thing you do before you ask, and became a step inside the model's own loop. You give it the question. It decides what to look up, runs the search, reads what came back, and searches again if the answer is not there yet.
Two details from this month make the shift concrete. GPT-6 Astra's documentation lists web search, file search and the Model Context Protocol as supported tools. It also lists tool search, which lets the model look through its own available tools rather than being handed a fixed list. Anthropic describe Opus 5.5 as a reliable researcher that found hard to locate sources in testing. The context window on Astra is now 1,050,000 tokens, with a knowledge cutoff of 30 April 2026. That gap between what a model knows and what is true is exactly the gap search exists to close.
The plumbing underneath this is worth knowing about, because it is the part that will outlive any particular model. MCP is the connector standard Anthropic published in November 2024, adopted by OpenAI in March 2025, and donated to the Agentic AI Foundation under the Linux Foundation on 9 December 2025, co-founded with Block and OpenAI. Google, Microsoft, Amazon Web Services, Cloudflare and Bloomberg all backed it. Anthropic reported 97 million monthly SDK downloads and 10,000 active servers at that point, which is the sort of number that makes a standard hard to reverse. It is now how an assistant reaches your systems rather than just the open web, which is the direction we wrote about in MCP and the connected office.
Anthropic also published figures on the unglamorous problem underneath long searches, which is that agents run out of room. Their context management work reports a 39 per cent improvement when a memory tool is combined with context editing. A 100 turn web search evaluation completed at 84 per cent lower token consumption, instead of failing outright.
There is a catch, and OpenAI name it themselves. Among the alignment evaluations published with Sol and Luna is one called broken search: the case where the tool fails and the model carries on as though it did not. A model that searches for you can be confidently wrong about what it found. Anything that matters still needs a source you can open and read.
Agents got further, and still need checking
More than last year, and less than the launch pages imply.
The genuine examples are striking. The 680,000 line migration above is one, work Anthropic estimate would have taken an engineering team weeks. OpenAI's Astra announcement shows it laying out a printed circuit board in KiCad, filling in a tax return and updating records in a CRM. It also ran front end quality checks on a site it had just built. On OSWorld 2.0, which tests long computer use workflows, Astra reached higher performance in about 47 per cent less time per task than the previous generation.
Now the number that belongs in the same paragraph. On Agents' Last Exam, which tests agents on real professional tasks across 55 sub-industries, Astra scored 59.3 per cent against Claude Opus 5 at 55.5 per cent. Fifty nine per cent is a real capability. It also means four attempts in ten came back wrong.
That is not an argument against using agents. It is an argument about where you put them. Work that is checked anyway, or cheap to redo, or where being wrong is visible immediately, is good work for an agent. Work where a wrong answer travels quietly into a valuation, a programme or a client email needs a person between the agent and the consequence. The guardrails have not changed and we set them out in full in agentic AI and the guardrails that make it safe.
The benchmark tables do not line up
Not as a league table, and it is worth understanding why before anyone builds a business case on one.
Take OSWorld 2.0, which both companies cite. Anthropic report Opus 5.5 at 81.8 per cent partial reward. OpenAI report Astra at 72.6 per cent in a latency simulation, and Sol at 60.5 per cent on the offline set from a specific August release. Those are different sets, different effort settings and different reward definitions. Neither company is being dishonest and the numbers still cannot be lined up.
OpenAI's own footnote is the most instructive thing on either page. They note that their cost figure for a competitor understates it. It leaves out the cost of fallbacks to another model, which happened on about 40 per cent of tasks. Read that as a warning about your own measurements rather than a point about theirs. The cost of an agent includes the retries, the fallbacks and the runs you threw away. A benchmark table will not tell you what those look like on your work.
The practical version of this is boring and it works. Take ten real tasks from your own week, run them on two models, and count how many came back usable and what each attempt cost. That afternoon is worth more than every comparison chart published this month. The framework we use for it is in what actually matters when a frontier model launches, and the model by model view is in which AI model for your team.
Four things worth doing this quarter
In this order.
- Move the routine work down a tier. If you are paying flagship rates to summarise documents, extract fields or answer known questions, you are paying for reasoning you are not using. That work belongs on the cheapest model that passes your own test, and the cheapest tier is now a tenth of what it cost in July.
- Measure cost per completed task. Not per million tokens. Include the retries and the runs you discarded, because those are most of the difference between a pilot that looks cheap and a bill that is not.
- Fix the caching before you change the model. Put the stable part of the prompt first, keep it stable, and check the cache is actually being hit. A 90 per cent discount on the reused part is a bigger saving than most model switches.
- Decide the permission boundary in writing. What the agent may read, what it may change, what needs a human yes, and where the record of what it did is kept. This is the part that does not expire when the next model ships, and the part almost nobody does first.
What are people asking us?
Do I need to change the model I am using?
Probably, but not to the biggest one. The useful change this month is at the cheaper end. GPT-6 Sol is half the price of the model it replaces. It also scores above Claude Opus 5 on OpenAI's own business workflow test, at a fraction of the cost per task. If you are paying flagship rates for work that is mostly reading, extracting and filing, you are paying for reasoning you are not using.
Is a cheaper model less accurate?
Not reliably, and that is the thing that changed. On OpenAI's internal factuality evaluation, GPT-6 Sol makes about half as many mistakes as the model it replaces. Accuracy and price used to move together. This month they did not, which is why the sensible move is to test rather than assume.
What is agentic search, in plain terms?
The model decides what to look up, runs the search itself, reads the results, and searches again if the answer is not there. You ask a question once instead of running six searches and reading twelve pages. The catch is that it can be confident about something it half found, so anything that matters still needs a source you can open.
Can an agent be left to run on its own?
Not on anything with consequences. On Agents' Last Exam, which tests agents on real professional tasks, the best score in OpenAI's comparison was 59.3 per cent. That is genuinely useful and it also means roughly four attempts in ten are wrong. Treat it as work that gets checked, and put an approval step in front of anything that spends money, sends a message or changes a record.
What does this actually cost to run?
Cost per completed task, not the headline rate. A cheap model that loops twenty times can cost more than an expensive one that gets it right first time. Caching is the lever most people miss. Cached input on GPT-6 is discounted by 90 per cent, and Opus 5.5 reads cached tokens at a fifth of the price Opus 5 did. Structure the prompt so the fixed part is reused and the bill drops without touching quality.
Is it worth waiting for the next release?
No, because there is always one about six weeks out. The models are now good enough that the limiting factor is your own process. That means what the agent is allowed to touch, where a human signs off, and whether the result is written down anywhere. That work does not go stale when the next model ships.
The question to ask before the next one ships
There will be another release in about six weeks, and the honest position is that the models are already ahead of most businesses' ability to use them. The limit is not the model. It is whether the work you would point it at is written down anywhere it can reach.
If you have a process eating hours every week and you suspect this month changed the maths on it, that is worth twenty minutes.
Sources
Every figure and price in this edition was read from the page linked below on 28 September 2026. Where a number could not be confirmed at its publisher, it is not in the edition.
- OpenAIIntroducing GPT-6 Sol and Luna
- AnthropicIntroducing Claude Opus 5.5
- OpenAIGPT-6 Astra
- OpenAIGPT-6 Astra model documentation
- AnthropicContext management
- AnthropicDonating the Model Context Protocol and establishing the Agentic AI Foundation
Followed this week
Where the week was picked up. Every figure and price above was then checked at the publisher and drawn from that source, so no chart here is traced off a video.
- Matthew BermanGPT-6 SOL AND LUNA ARE OUT!!!
- Matt WolfeAI News: Opus 5.5, GPT-6 Sol, Jev, Muse and More!
- Matt WolfeClaude Opus 5.5 Didn’t Need to Go This Hard