I don't see how anyone can be using Claude with prices like this, it's pretty incredible what the OpenAI team is doing, w.r.t model quality and pricing.
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
These don’t necessarily reflect actual costs, OpenAI is not profitable and nowhere near. They’ve lost their market lead and Sam may feel they need to get it back with any means necessary.
GPT would charge more if they could. Both companies need way way more revenue. GPT simply made a calculation that they can earn more money by charging less than their competitors.
Any business would charge more if they could. Jevon's paradox would mean that they can make more money by charging less because demand is going to keep growing.
Ya, are LLM's not a great example of Jevon's paradox? I don't think Jevon's needs all else being equal. The paradox being that we should be able to use things less because they are more efficient, when instead they get used more.
Surely, a large part of the increase of the demand in LLMs is in their intelligence, but to hit the demand models needed to be made more efficient, and labs found that more efficient models, still demanded more usage.
They're cutting prices because they want to cannabalize the market for people using models like deepseek via API as well as people paying for anthropic subs.
When they cut prices on luna the first time around they took (literally) millions of users from anthropic.
Just because a provider is charging less, doesn't mean their cost went down. This is probably especially true with the big players that are trying to stay competitive.
A lot of what I'm doing has pretty expensive build/testing processes between iterations - even on a 40 core machine - so I'm not burning tokens 24/7 like some people may.
I'd guess I'm probably spending >50% of the time running tests & build processes & tooling and the remainder is purely burning tokens.
I also have some internal tooling (that I will hopefully open source soon) that makes LLMs substantially more correct (thus more efficient) - so there's that, too.
GPT-5.6-Sol, GPT-5.6-Terra, and GPT-5.6-Luna were released in July of 2026.
The first release from the GPT-6 series was GPT-6-Astra. GPT-6-Astra happened on around September 3, 2026, and the previously-mentioned GPT-5.6-* widgets remained available.
Today, September 22, 2026, we now also have GPT-6-Sol and GPT-6-Luna added into the mix.
As I write this, all of the model identifiers I've mentioned are available to select for use within Codex.
and anthropic won't? or any other inference provider? Running your own inference either locally or remotely are probably the only ways to make sure that doesn't happen.
And why is that bad? As your brain gets older, it will not remain so clever, so you'll be grateful for an AI that thinks like you do when it comes to your line of work, failing which the quality of your output could recede like your hairline.
You're building your livelihood/workflows on a set of inputs that you have no idea what they actually cost or how reliable they'll be when the VC cash stops flowing. If you're OK with that, do your thing but it seems a little foolish to me.
I wouldn't bet on hardware flooding the market. I bet the machines running in the data centers don't use traditional PCIe connectors and cards. Maybe somebody could pull the chips and put them on standardized PCIe cards, but that is not a given.
V100s are three generations behind current and missing many of the features that modern inference benefits from, but they are the cheapest way to get a 32GB gpu.
It seems silly to say we have no idea when we actually do, though. We know how much hardware costs, we know how to reliably run a webservice that hits an API hosted on a machine with a GPU, we know how to operate these things at scale outside of OpenAI and Anthropic (not Nvidia). VC money can be patient, Uber's profitable, yeah $1 Uber rides got us hooked and they're running the same playbook. Unfortunately the convenience is worth paying for, so it seems dumb to think we can control the beast or ignore it, or get everyone to agree to hold back.
Is there a world where OpenAI starts charging $2,000/month for what we previously were paying $20 for? What are we going to do? AWS could totally jack up the prices for EC2 instances as well, but we've come to rely on that as well.
Push comes to shove, OpenAI could go out of business tomorrow and I could pick up roughly where I left off for $25k, which is the cost to serve GLM 5.3 Flash on four Nvidia GB10s. Granted, if OpenAI et al go kaput all at the same time, I could probably get a whole lot more compute for a whole lot less money.
huh? i use the plans because they're cheap and i get strong models, but i could go back to deepseek flash on commodity api pricing and be just fine
I'm fairly sure most open weight model providers are serving them at a sustainable price - and I've used them enough to know that I could live with them if the big boys did a rug pull.
Not that anthropic models are very good at this, but due to the changes in tokenizers and thinking tokens: cost per token is not as helpful anymore as cost / task.
As someone who has used Claude Code and Codex the prices don't matter in the same way but I found that I burned through my usage way faster on Codex even though I regularly hear that the Codex plans go further. That was not my experience and the intelligence was comparable to what I was getting in Claude.
If these price changes mean that coding plans have effectively more usage then that's great, but Codex is surviving on resets from my own experience using it. I was glad to go back to Claude.
I mainly use Codex/Sol to review my plans drafted by Fable. But beyond that, Astra blows through usage limits too fast to be a daily driver and writes weird code despite what my "house style" is, and Codex is behind Claude Code in terms of critical features like seeing what's going on in subagents.
The parent + subagent workflow has become critical for keeping the reasoning agent (parent) context-lean while also letting me chat to the main agent while work is getting done.
My main process is to use Fable to reason and then spawn Opus subagents, and I get amazing results, and I'm always looking into what the subagents are doing.
- A model (like other GPT ones) that hides thinking traces and thinking summaries, which infuriates me
I've been in the Claude camp for a while, but the way it writes has left me with a a brick for a brain and wanted to see if Astra was as good as they say. Well, I can't know, because in the time it takes for it to actually build anything useful, I've moved to other ideas.
Unbearably, annoyingly slow. I keep thinking I must be doing something wrong.
You're right, it's probably quite unfair of me to say it eats lots of tokens when I am paying double for claude than codex and complaining about tokens.
The rest still stands, though.
But if I've learned anything is that in a 2 months I might have completely turned around, who knows
The thing is that this thing is constantly compacting.... I get 1M context with Claude and ~256k with Astra. Even if the compaction loses much less information on OAI's side, it takes so long it's barely any use for me...
I've tried High and Max. They have produced decent results, but they're so slow.... I will try to lower it a bit and see the difference, but it's a delicate balance: I don't want to waste literal hours on the incorrect reasoning level to only then have to spend those hours and tokens to do it right.
At this very moment, Astra has been working for 1h15m on a task. At this rate I genuinely expect it to take about 10 hours. I feel like claude would do it in at least a third of that. Let's see if the quality justifies the slowness (it better)
If you want a mix between both codex/claude code/pi for leveraging different models and harnesses, you can give ai-devkit agent orchestration a try
Let's look at open-weights models with 3T size: <a href="https://inferencex.semianalysis.com/run/kimi-k3-on-b200" rel="nofollow">https://inferencex.semianalysis.com/run/kimi-k3-on-b200
This suggests inference margins in the ballpark of 98% if we assume 5.6 Sol is about as efficient to serve as Kimi K3.
We also do not know what efficiency improvements have been made with GPT 6 Sol and Luna.
There is some speculation that 6 Sol could be a smaller model comparable in size to 5.6 Terra, and that this is why the improvement in intelligence is modest over 5.6 Sol.
This would line up with a faster serving speed and benchmarks that show a small improvement in coding tasks with regressions in knowledge tasks.
The major difference being the 1M token context window. Once you exceed 272K input tokens, Codex Sol is roughly the same price as Opus; and Astra similar to Fable.
>I don't see how anyone can be using Claude with prices like this
One potential deciding point is that Claude still has a $200/mo 20x plan, where, since Sept 11, OpenAI does not and has no ETA for the return.
I downgraded my OpenAI plan 2 months ago to the $100/mo, but my usage has gone way up, but now I can no longer upgrade to the $200/mo plan ("This option is temporarily unavailable"). Thankfully I have 2 usage resets available, but I'll probably be switching back to Claude; I was super happy with Astra but I'm burning through tokens and have 4 days before my next reset.
Grok is offering a competitive product - not the absolute best but among the top three. They are doing for $2 and 6/mio. So maybe they see it as taking the economic opportunity while their model lags slightly behind. OpenAi follows. I can't say if their models are economically better or they are taking the loss but they can still pull it. SpaceXAI has an interesting path forward. They are not out.
pookieinc · · focus · HN ↗
Input
Output
Price reduction
GPT‑6 Sol vs. GPT‑5.6 Sol
$4 → $2
$20 → $10
50% cheaper
GPT‑6 Luna vs. GPT‑5.6 Luna
$0.20 → $0.10
$1.20 → $0.50
50% cheaper
thereitgoes456 · · focus · HN ↗
giancarlostoro · · focus · HN ↗
andybak · · focus · HN ↗
singingtoday · · focus · HN ↗
Readerium · · focus · HN ↗
Should be B vs A correct?
Else it's confusing
bitmasher9 · · focus · HN ↗
wyre · · focus · HN ↗
atq2119 · · focus · HN ↗
wyre · · focus · HN ↗
Surely, a large part of the increase of the demand in LLMs is in their intelligence, but to hit the demand models needed to be made more efficient, and labs found that more efficient models, still demanded more usage.
blovescoffee · · focus · HN ↗
blubber · · focus · HN ↗
vanuatu · · focus · HN ↗
Shekelphile · · focus · HN ↗
When they cut prices on luna the first time around they took (literally) millions of users from anthropic.
mchusma · · focus · HN ↗
Its a great release, I will use both heavily.
etothet · · focus · HN ↗
esafak · · focus · HN ↗
etothet · · focus · HN ↗
mfiguiere · · focus · HN ↗
<a href="https://developers.openai.com/api/docs/pricing?latest-pricing=batch" rel="nofollow">https://developers.openai.com/api/docs/pricing?latest-pricin...
onlyrealcuzzo · · focus · HN ↗
This is great, but practically, I'm not going to start working on more side projects.
Perhaps in another 6-12 months I'll be fine to drop down to $20/m instead of $200.
charliegoforit · · focus · HN ↗
wyre · · focus · HN ↗
onlyrealcuzzo · · focus · HN ↗
A lot of what I'm doing has pretty expensive build/testing processes between iterations - even on a 40 core machine - so I'm not burning tokens 24/7 like some people may.
I'd guess I'm probably spending >50% of the time running tests & build processes & tooling and the remainder is purely burning tokens.
I also have some internal tooling (that I will hopefully open source soon) that makes LLMs substantially more correct (thus more efficient) - so there's that, too.
szundi · · focus · HN ↗
[dead]
adam_arthur · · focus · HN ↗
There are a ton of use cases that open up with cheaper models.
E.g. extensive security scanning on every PR, quality scans, adversarial reviews etc
onlyrealcuzzo · · focus · HN ↗
shmoil · · focus · HN ↗
>> $4 → $2
>> $20 → $10
Do you mean 100% more expensive? GPT 6 is 100% more expensive than 5.6 per your post.
blovescoffee · · focus · HN ↗
s3p · · focus · HN ↗
yzydserd · · focus · HN ↗
jameshart · · focus · HN ↗
edf13 · · focus · HN ↗
dyauspitr · · focus · HN ↗
Readerium · · focus · HN ↗
dyauspitr · · focus · HN ↗
MaKey · · focus · HN ↗
Readerium · · focus · HN ↗
Performance increases both with larger model (Luna vs Sol)
And with more reasoning (low vs xhigh)
hersko · · focus · HN ↗
ssl-3 · · focus · HN ↗
GPT-5.6-Sol, GPT-5.6-Terra, and GPT-5.6-Luna were released in July of 2026.
The first release from the GPT-6 series was GPT-6-Astra. GPT-6-Astra happened on around September 3, 2026, and the previously-mentioned GPT-5.6-* widgets remained available.
Today, September 22, 2026, we now also have GPT-6-Sol and GPT-6-Luna added into the mix.
As I write this, all of the model identifiers I've mentioned are available to select for use within Codex.
LZ_Khan · · focus · HN ↗
AustinDev · · focus · HN ↗
solenoid0937 · · focus · HN ↗
vanuatu · · focus · HN ↗
nradov · · focus · HN ↗
OutOfHere · · focus · HN ↗
redanddead · · focus · HN ↗
Highly subjective take
What kind of work do you do, out of curiosity
copperx · · focus · HN ↗
OutOfHere · · focus · HN ↗
LZ_Khan · · focus · HN ↗
trentor · · focus · HN ↗
artursapek · · focus · HN ↗
trentor · · focus · HN ↗
an0malous · · focus · HN ↗
solenoid0937 · · focus · HN ↗
blovescoffee · · focus · HN ↗
selectodude · · focus · HN ↗
If they’re subsidizing my usage, that’s great.
infinitezest · · focus · HN ↗
derac · · focus · HN ↗
ssl-3 · · focus · HN ↗
foepys · · focus · HN ↗
Leynos · · focus · HN ↗
External example: <a href="https://ebay.io/m/lV8UsD" rel="nofollow">https://ebay.io/m/lV8UsD
Internal example: <a href="https://ebay.io/m/z1ygRU" rel="nofollow">https://ebay.io/m/z1ygRU
V100s are three generations behind current and missing many of the features that modern inference benefits from, but they are the cheapest way to get a 32GB gpu.
fragmede · · focus · HN ↗
Is there a world where OpenAI starts charging $2,000/month for what we previously were paying $20 for? What are we going to do? AWS could totally jack up the prices for EC2 instances as well, but we've come to rely on that as well.
selectodude · · focus · HN ↗
slopinthebag · · focus · HN ↗
goosejuice · · focus · HN ↗
andybak · · focus · HN ↗
minimaxir · · focus · HN ↗
persedes · · focus · HN ↗
joshstrange · · focus · HN ↗
If these price changes mean that coding plans have effectively more usage then that's great, but Codex is surviving on resets from my own experience using it. I was glad to go back to Claude.
hombre_fatal · · focus · HN ↗
The parent + subagent workflow has become critical for keeping the reasoning agent (parent) context-lean while also letting me chat to the main agent while work is getting done.
My main process is to use Fable to reason and then spawn Opus subagents, and I get amazing results, and I'm always looking into what the subagents are doing.
jorl17 · · focus · HN ↗
- Unbearably slow
- A token eating machine like no other
- Constantly compacting
- A model (like other GPT ones) that hides thinking traces and thinking summaries, which infuriates me
I've been in the Claude camp for a while, but the way it writes has left me with a a brick for a brain and wanted to see if Astra was as good as they say. Well, I can't know, because in the time it takes for it to actually build anything useful, I've moved to other ideas.
Unbearably, annoyingly slow. I keep thinking I must be doing something wrong.
user43928 · · focus · HN ↗
However, it is not a 'token eating machine'. In fact it uses a third of the output tokens of Opus 5.5, Fable 5.1, or Opus 5.
17k for Astra xhigh vs 61-66k.
jorl17 · · focus · HN ↗
The rest still stands, though.
But if I've learned anything is that in a 2 months I might have completely turned around, who knows
thatguymike · · focus · HN ↗
jorl17 · · focus · HN ↗
I've tried High and Max. They have produced decent results, but they're so slow.... I will try to lower it a bit and see the difference, but it's a delicate balance: I don't want to waste literal hours on the incorrect reasoning level to only then have to spend those hours and tokens to do it right.
At this very moment, Astra has been working for 1h15m on a task. At this rate I genuinely expect it to take about 10 hours. I feel like claude would do it in at least a third of that. Let's see if the quality justifies the slowness (it better)
hoangnnguyen · · focus · HN ↗
ignoramous · · focus · HN ↗
Cache read/write decrease by 50% or similar? That's where most (95%+) of the cost is for agentic coding workloads.
minimaxir · · focus · HN ↗
<a href="https://developers.openai.com/api/docs/pricing" rel="nofollow">https://developers.openai.com/api/docs/pricing
minimaxir · · focus · HN ↗
user43928 · · focus · HN ↗
Let's look at open-weights models with 3T size: <a href="https://inferencex.semianalysis.com/run/kimi-k3-on-b200" rel="nofollow">https://inferencex.semianalysis.com/run/kimi-k3-on-b200
This suggests inference margins in the ballpark of 98% if we assume 5.6 Sol is about as efficient to serve as Kimi K3.
We also do not know what efficiency improvements have been made with GPT 6 Sol and Luna.
There is some speculation that 6 Sol could be a smaller model comparable in size to 5.6 Terra, and that this is why the improvement in intelligence is modest over 5.6 Sol.
This would line up with a faster serving speed and benchmarks that show a small improvement in coding tasks with regressions in knowledge tasks.
baalimago · · focus · HN ↗
We don't know how much they are bleeding financially, it might just be a front
rgbrenner · · focus · HN ↗
linsomniac · · focus · HN ↗
One potential deciding point is that Claude still has a $200/mo 20x plan, where, since Sept 11, OpenAI does not and has no ETA for the return.
I downgraded my OpenAI plan 2 months ago to the $100/mo, but my usage has gone way up, but now I can no longer upgrade to the $200/mo plan ("This option is temporarily unavailable"). Thankfully I have 2 usage resets available, but I'll probably be switching back to Claude; I was super happy with Astra but I'm burning through tokens and have 4 days before my next reset.
sick_of_slop · · focus · HN ↗
[dead]
freeandclear · · focus · HN ↗
thefourthchime · · focus · HN ↗