GPT-6 Sol and Luna
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
GPT-6 Sol and Luna
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
hehimself · · focus · HN ↗
madduci · · focus · HN ↗
blovescoffee · · focus · HN ↗
eloisant · · focus · HN ↗
system2 · · focus · HN ↗
Also Mimo 2.6 is roughly 30% cheaper. Without batch.
Havoc · · focus · HN ↗
system2 · · focus · HN ↗
Havoc · · focus · HN ↗
Or use the failure to get a response like you say
system2 · · focus · HN ↗
beardsciences · · focus · HN ↗
jstummbillig · · focus · HN ↗
jrflo · · focus · HN ↗
Cu3PO42 · · focus · HN ↗
yreg · · focus · HN ↗
manmal · · focus · HN ↗
manmal · · focus · HN ↗
iyonn · · focus · HN ↗
c0rruptbytes · · focus · HN ↗
motoboi · · focus · HN ↗
whazor · · focus · HN ↗
potwinkle · · focus · HN ↗
Readerium · · focus · HN ↗
hehimself · · focus · HN ↗
Readerium · · focus · HN ↗
blovescoffee · · focus · HN ↗
pinkgolem · · focus · HN ↗
droidjj · · focus · HN ↗
scrlk · · focus · HN ↗
simianwords · · focus · HN ↗
[dead]
nickandbro · · focus · HN ↗
physicallyIllfr · · focus · HN ↗
Does anyone care about code quality anymore?
blovescoffee · · focus · HN ↗
physicallyIllfr · · focus · HN ↗
This will make me more valuable in the future when everyone has lost the ability to do anything on their own.
jpadkins · · focus · HN ↗
I respect that you want to learn how things are done, that is a great trait. But once you learn how its done, you should use the tools to free up cognitive load for more difficult tasks.
fragmede · · focus · HN ↗
minimaxir · · focus · HN ↗
physicallyIllfr · · focus · HN ↗
thimabi · · focus · HN ↗
jasbury · · focus · HN ↗
pookieinc · · focus · HN ↗
Input
Output
Price reduction
GPT‑6 Sol vs. GPT‑5.6 Sol
$4 → $2
$20 → $10
50% cheaper
GPT‑6 Luna vs. GPT‑5.6 Luna
$0.20 → $0.10
$1.20 → $0.50
50% cheaper
thereitgoes456 · · focus · HN ↗
giancarlostoro · · focus · HN ↗
andybak · · focus · HN ↗
singingtoday · · focus · HN ↗
Readerium · · focus · HN ↗
Should be B vs A correct?
Else it's confusing
bitmasher9 · · focus · HN ↗
wyre · · focus · HN ↗
atq2119 · · focus · HN ↗
wyre · · focus · HN ↗
Surely, a large part of the increase of the demand in LLMs is in their intelligence, but to hit the demand models needed to be made more efficient, and labs found that more efficient models, still demanded more usage.
blovescoffee · · focus · HN ↗
blubber · · focus · HN ↗
vanuatu · · focus · HN ↗
Shekelphile · · focus · HN ↗
When they cut prices on luna the first time around they took (literally) millions of users from anthropic.
mchusma · · focus · HN ↗
Its a great release, I will use both heavily.
etothet · · focus · HN ↗
esafak · · focus · HN ↗
etothet · · focus · HN ↗
mfiguiere · · focus · HN ↗
<a href="https://developers.openai.com/api/docs/pricing?latest-pricing=batch" rel="nofollow">https://developers.openai.com/api/docs/pricing?latest-pricin...
onlyrealcuzzo · · focus · HN ↗
This is great, but practically, I'm not going to start working on more side projects.
Perhaps in another 6-12 months I'll be fine to drop down to $20/m instead of $200.
charliegoforit · · focus · HN ↗
wyre · · focus · HN ↗
onlyrealcuzzo · · focus · HN ↗
A lot of what I'm doing has pretty expensive build/testing processes between iterations - even on a 40 core machine - so I'm not burning tokens 24/7 like some people may.
I'd guess I'm probably spending >50% of the time running tests & build processes & tooling and the remainder is purely burning tokens.
I also have some internal tooling (that I will hopefully open source soon) that makes LLMs substantially more correct (thus more efficient) - so there's that, too.
szundi · · focus · HN ↗
[dead]
adam_arthur · · focus · HN ↗
There are a ton of use cases that open up with cheaper models.
E.g. extensive security scanning on every PR, quality scans etc
onlyrealcuzzo · · focus · HN ↗
shmoil · · focus · HN ↗
>> $4 → $2
>> $20 → $10
Do you mean 100% more expensive? GPT 6 is 100% more expensive than 5.6 per your post.
blovescoffee · · focus · HN ↗
s3p · · focus · HN ↗
yzydserd · · focus · HN ↗
jameshart · · focus · HN ↗
edf13 · · focus · HN ↗
dyauspitr · · focus · HN ↗
Readerium · · focus · HN ↗
dyauspitr · · focus · HN ↗
MaKey · · focus · HN ↗
Readerium · · focus · HN ↗
Performance increases both with larger model (Luna vs Sol)
And with more reasoning (low vs xhigh)
hersko · · focus · HN ↗
ssl-3 · · focus · HN ↗
GPT-5.6-Sol, GPT-5.6-Terra, and GPT-5.6-Luna were released in July of 2026.
The first release from the GPT-6 series was GPT-6-Astra. GPT-6-Astra happened on around September 3, 2026, and the previously-mentioned GPT-5.6-* widgets remained available.
Today, September 22, 2026, we now also have GPT-6-Sol and GPT-6-Luna added into the mix.
As I write this, all of the model identifiers I've mentioned are available to select for use within Codex.
LZ_Khan · · focus · HN ↗
AustinDev · · focus · HN ↗
solenoid0937 · · focus · HN ↗
vanuatu · · focus · HN ↗
nradov · · focus · HN ↗
OutOfHere · · focus · HN ↗
redanddead · · focus · HN ↗
Highly subjective take
What kind of work do you do, out of curiosity
copperx · · focus · HN ↗
OutOfHere · · focus · HN ↗
LZ_Khan · · focus · HN ↗
trentor · · focus · HN ↗
artursapek · · focus · HN ↗
trentor · · focus · HN ↗
an0malous · · focus · HN ↗
solenoid0937 · · focus · HN ↗
blovescoffee · · focus · HN ↗
selectodude · · focus · HN ↗
If they’re subsidizing my usage, that’s great.
infinitezest · · focus · HN ↗
derac · · focus · HN ↗
ssl-3 · · focus · HN ↗
foepys · · focus · HN ↗
Leynos · · focus · HN ↗
External example: <a href="https://ebay.io/m/lV8UsD" rel="nofollow">https://ebay.io/m/lV8UsD
Internal example: <a href="https://ebay.io/m/z1ygRU" rel="nofollow">https://ebay.io/m/z1ygRU
V100s are three generations behind current and missing many of the features that modern inference benefits from, but they are the cheapest way to get a 32GB gpu.
fragmede · · focus · HN ↗
Is there a world where OpenAI starts charging $2,000/month for what we previously were paying $20 for? What are we going to do? AWS could totally jack up the prices for EC2 instances as well, but we've come to rely on that as well.
selectodude · · focus · HN ↗
slopinthebag · · focus · HN ↗
goosejuice · · focus · HN ↗
andybak · · focus · HN ↗
minimaxir · · focus · HN ↗
persedes · · focus · HN ↗
joshstrange · · focus · HN ↗
If these price changes mean that coding plans have effectively more usage then that's great, but Codex is surviving on resets from my own experience using it. I was glad to go back to Claude.
hombre_fatal · · focus · HN ↗
I use Fable to spawn Opus subagents and have amazing results, and I'm always looking into what the subagents are doing.
jorl17 · · focus · HN ↗
- Unbearably slow - A token eating machine like no other - Constantly compacting - A model (like other GPT ones) that hides thinking traces and thinking summaries, which infuriates me
I've been in the Claude camp for a while, but the way it writes has left me with a a brick for a brain and wanted to see if Astra was as good as they say. Well, I can't know, because in the time it takes for it to actually build anything useful, I've moved to other ideas.
Unbearably, annoyingly slow. I keep thinking I must be doing something wrong.
user43928 · · focus · HN ↗
However, it is not a 'token eating machine'. In fact it uses a third of the output tokens of Opus 5.5, Fable 5.1, or Opus 5.
17k for Astra xhigh vs 61-66k.
jorl17 · · focus · HN ↗
The rest still stands, though.
But if I've learned anything is that in a 2 months I might have completely turned around, who knows
thatguymike · · focus · HN ↗
jorl17 · · focus · HN ↗
I've tried High and Max. They have produced decent results, but they're so slow.... I will try to lower it a bit and see the difference, but it's a delicate balance: I don't want to waste literal hours on the incorrect reasoning level to only then have to spend those hours and tokens to do it right.
At this very moment, Astra has been working for 1h15m on a task. At this rate I genuinely expect it to take about 10 hours. I feel like claude would do it in at least a third of that. Let's see if the quality justifies the slowness (it better)
hoangnnguyen · · focus · HN ↗
ignoramous · · focus · HN ↗
minimaxir · · focus · HN ↗
minimaxir · · focus · HN ↗
user43928 · · focus · HN ↗
Let's look at open-weights models with 3T size: <a href="https://inferencex.semianalysis.com/run/kimi-k3-on-b200" rel="nofollow">https://inferencex.semianalysis.com/run/kimi-k3-on-b200
This suggests inference margins in the ballpark of 98% if we assume 5.6 Sol is about as efficient to serve as Kimi K3.
baalimago · · focus · HN ↗
We don't know how much they are bleeding financially, it might just be a front
rgbrenner · · focus · HN ↗
linsomniac · · focus · HN ↗
One potential deciding point is that Claude still has a $200/mo 20x plan, where, since Sept 11, OpenAI does not and has no ETA for the return.
I downgraded my OpenAI plan 2 months ago to the $100/mo, but my usage has gone way up, but now I can no longer upgrade to the $200/mo plan ("This option is temporarily unavailable"). Thankfully I have 2 usage resets available, but I'll probably be switching back to Claude; I was super happy with Astra but I'm burning through tokens and have 4 days before my next reset.
sick_of_slop · · focus · HN ↗
[dead]
freeandclear · · focus · HN ↗
thefourthchime · · focus · HN ↗
mchusma · · focus · HN ↗
sfkgtbor · · focus · HN ↗
Readerium · · focus · HN ↗
Wierd!!
samuelknight · · focus · HN ↗
MrBuddyCasino · · focus · HN ↗
o_m · · focus · HN ↗
minimaxir · · focus · HN ↗
MangoCoffee · · focus · HN ↗
Astra is the best. Luna is cheapest then it seems like Sol is the middle child like Terra.
scrlk · · focus · HN ↗
cesarvarela · · focus · HN ↗
petesergeant · · focus · HN ↗
afro88 · · focus · HN ↗
cesarvarela · · focus · HN ↗
recitedropper · · focus · HN ↗
wahnfrieden · · focus · HN ↗
dominotw · · focus · HN ↗
rtaylorgarlock · · focus · HN ↗
ronsor · · focus · HN ↗
But the reason people say "Claude can't compete" is because Claude Opus has been going downhill since 4.7, and many have found Opus 5 intolerable. Fable is much better, but also much more expensive than OpenAI's offerings.
CaptWorld · · focus · HN ↗
droidjj · · focus · HN ↗
mydreamof · · focus · HN ↗
saadn92 · · focus · HN ↗
civvv · · focus · HN ↗
ZeWaka · · focus · HN ↗
LeBit · · focus · HN ↗
qoez · · focus · HN ↗
john_strinlai · · focus · HN ↗
LeBit · · focus · HN ↗
adrianwaj · · focus · HN ↗
<a href="https://hackertrain.future-secured.com/?q=great" rel="nofollow">https://hackertrain.future-secured.com/?q=great
One thing that's changed over time is a lot more usage of "scare quotes" especially now (Gemini agrees <a href="https://share.gemini.google/UiUf0jttZLVD" rel="nofollow">https://share.gemini.google/UiUf0jttZLVD ). It's an interesting phenomenon, Abloh started using quotation marks consistently in his fashion branding since 2012. <a href="https://blakecrosley.com/blog/design-philosophy-virgil-abloh" rel="nofollow">https://blakecrosley.com/blog/design-philosophy-virgil-abloh
Next step would be to attach an AI to all the /bestcomments.. if someone needs help doing that I'm here. Really, that's a task for the mods.
monkeydust · · focus · HN ↗
Madmallard · · focus · HN ↗
minimaxir · · focus · HN ↗
Madmallard · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
woah · · focus · HN ↗
sidrag22 · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
Asking cuz I don't think I'm a bot, I legitimately prefer the GPT models to Anthropic's, don't like Anthropic's customer service/reliability story at all, and I welcome a massive price reduction. Seems like something I should be happy to get.
theanonymousone · · focus · HN ↗
Eldodi · · focus · HN ↗
badatnames · · focus · HN ↗
chaos_emergent · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
Then they do a new model launch, issue quota resets all around, and it's a party for 2-3 weeks before things return to normal.
FergusArgyll · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
badatnames · · focus · HN ↗
wyre · · focus · HN ↗
badatnames · · focus · HN ↗
OpenAI/Anthropic meanwhile feel a bit like they're hoping to sell iPhones in a market about to be flooded by $20 flip phones, with almost no channel of their own to do it. And for whatever mad reason OpenAI are now signalling they will attempt to compete on price with flip phones despite their cost of labour, energy, and just about everything else being far higher
wyre · · focus · HN ↗
I don't see your metaphor to iphones and flip phones. This new Luna model is cheaper than deepseek 4.1 flash, except for cache reads. OpenAI having to compete with China is a much larger economic-political issue that is far larger than just our AI labs.
PP9866 · · focus · HN ↗
staticman2 · · focus · HN ↗
kibae · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
Which... fine, I'll take that.
cmrdporcupine · · focus · HN ↗
Ninjinka · · focus · HN ↗
Readerium · · focus · HN ↗
eyk19 · · focus · HN ↗
mrdependable · · focus · HN ↗
jrflo · · focus · HN ↗
m_fayer · · focus · HN ↗
redox99 · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
Astra was/is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague/lazy instructions ("I need to be able to test this on Windows, maybe a qemu VM or something? Shrug." ... 1 hour later "yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve & verify harness.").
And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.
But it also feels sloppier? Somehow. And too expensive to use.
We'll see how Sol 6 is.
jeffnash · · focus · HN ↗
Terra had the "workhorse" quality where it could do these changes in bulk and follow directions without being too 'smart' (but sloppy) as you described. Luna was a bit too dumb and would make sloppy mistakes; I see that more as a "run these tests and format the results" sort of model. Maybe 6 Luna will be better.
I also just reread your comment and realized the naming convention is still extremely confusing with respect to ordering of [Family]x[Model]x[Number].
m_fayer · · focus · HN ↗
jeffnash · · focus · HN ↗
fodkodrasz · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
When people spend their days interacting with machines that pretend to be human, they may then start treating real humans like machines.
yomismoaqui · · focus · HN ↗
buu700 · · focus · HN ↗
4b11b4 · · focus · HN ↗
mavsman · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
Sol 6 is a heaping pile of garbage. Just epic levels of slop. And r/codex etc is full of people noticing the same.
I've switched back to 5.6 Sol. What they're selling as Sol 6 is really what would have been Terra before, and it's awful.
Rapzid · · focus · HN ↗
Otherwise I'm using 5.6 Sol for actual plan execution and review..
amluto · · focus · HN ↗
(I have not done anything quantitative here. For one thing, OpenAI’s billing pages and the codex-rs frontend make it pathetically difficult to get any real data. Some day I should wire up a proxy to extract actual stats.)
NorthSouthNorth · · focus · HN ↗
jauntywundrkind · · focus · HN ↗
It wants things beyond what the mortals (us) know to reach for. It's not good at explaining itself, it doesn't show it's thinking. It's often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.
dannyw · · focus · HN ↗
The other explanation is just as part of ‘token efficiency’
throwuxiytayq · · focus · HN ↗
jauntywundrkind · · focus · HN ↗
that could help tackle half of the problems here. i do think the other 100% completionist part is something i'm more used to steering through with llm usage, have negotiated fora while, and that Astra is particularly an astronaut whose instincts are extremely strongly in the direction of foreseeing and outdesigning potential problems, that it is rarely going to pick a practical sensible clear path on it's own.
fnordpiglet · · focus · HN ↗
bryanhogan · · focus · HN ↗
My results with 5.6 Sol were quite similar to 6, although I haven't tested it that much.
nickreese · · focus · HN ↗
mcast · · focus · HN ↗
jdw64 · · focus · HN ↗
BowBun · · focus · HN ↗
bradly · · focus · HN ↗
Highlight and lowlight of my week was successfully convincing the OpenAI support chat robot to give me a refund for the month for my issues with 6 chewing threw my usage with no output.
AaronAPU · · focus · HN ↗
apitman · · focus · HN ↗
pyed · · focus · HN ↗
[dead]
simianwords · · focus · HN ↗
alansaber · · focus · HN ↗
4b11b4 · · focus · HN ↗
jmuguy · · focus · HN ↗
danabramov · · focus · HN ↗
If I leave Astra overnight, I'll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.
jijijijij · · focus · HN ↗
m_fayer · · focus · HN ↗
sinsterizme · · focus · HN ↗
Like you said, it felt very natural to work with. Opus 5 is way too slow and verbose for me, I find I get distracted and annoyed with it.
Opus 5.5 seems a LOT closer so far to what I liked about 5.6 Sol but we'll see
bredren · · focus · HN ↗
Except they were not for past few years as they misfired on the attempt to compete with vscode. That had a big impact on pycharm, which seemed starved for resources for so long. The company eventually declared a year of Django, but even that failed to really make an impact.
Arguably, Jetbrains had first insight into AI based code completion via rapid rise of the TabNine plugin but missed that opportunity also.
capital_guy · · focus · HN ↗
if GPT 6 Sol is just 5.6 at half the price it will be everything i really ever wanted.
manojlds · · focus · HN ↗
makeavish · · focus · HN ↗
joseda-hg · · focus · HN ↗
They usually reduce usage consumption in line with cost reductions (But not always 1:1)
flippingheck · · focus · HN ↗
How are people using 5.6 Sol? API pricing? Subscriptions?
I like because DeepSeek 4.1 Flash because I never experience quota issues, and it's still cheap and mostly good enough.
jrflo · · focus · HN ↗
flippingheck · · focus · HN ↗
4b11b4 · · focus · HN ↗
cedws · · focus · HN ↗
joduplessis · · focus · HN ↗
Imanari · · focus · HN ↗
ljm · · focus · HN ↗
But I wonder if that's intentional because it can keep computing while you are answering, so long as your steer aligns well enough with the direction it wants to go. Better than letting a cache go cold and burning compute on bringing it all back up.
meerita · · focus · HN ↗
slekker · · focus · HN ↗
system2 · · focus · HN ↗
Nvidia's top AI chip Rubin sells in 72-GPU racks for about $3.5–7.8M. A rack running Xiaomi's MiMo V2.6 Pro generates roughly 150–300B tokens a day, worth about $130–260k at Xiaomi's API price. That's a payback of the infrastructure in a few weeks in theory. After a few weeks or a month, the only cost is electricity, and whatever they make after that is pure profit.
OpenAI and Anthropic are practically scamming people with the token prices.
thimabi · · focus · HN ↗
wyre · · focus · HN ↗
desterothx · · focus · HN ↗
anthonyrstevens · · focus · HN ↗
simianparrot · · focus · HN ↗
sehw · · focus · HN ↗
msh · · focus · HN ↗
blovescoffee · · focus · HN ↗
OutOfHere · · focus · HN ↗
Tadpole9181 · · focus · HN ↗
Terra ended up just being an awkward middle ground that was not particularly suited for any workload.
OutOfHere · · focus · HN ↗
Tadpole9181 · · focus · HN ↗
msh · · focus · HN ↗
darklinear · · focus · HN ↗
Maybe for one-shotting large things Sol is better, but for prod code where I decompose into smaller tasks and read all the code I favored Terra.
Sol 5.6 was still king for architecture/research in my workflow, though.
smith7018 · · focus · HN ↗
yawnxyz · · focus · HN ↗
sandos · · focus · HN ↗
Funny thing is they very recently also set a real limit per-user/month, so why even limit the models because theyre "too expensive".
apitman · · focus · HN ↗
miohtama · · focus · HN ↗
devinprater · · focus · HN ↗
markerbrod · · focus · HN ↗
Edit: Yes, it applies also to subscriptions, source <a href="https://x.com/thsottiaux/status/2102463847714247142" rel="nofollow">https://x.com/thsottiaux/status/2102463847714247142
jrflo · · focus · HN ↗
fHr · · focus · HN ↗
yipinwong · · focus · HN ↗
Now GPT 6 Luna is even cheaper, and more intelligent, there is no going back... to SOL 5.6 for intelligent layer.
dmazin · · focus · HN ↗
I was hoping for a serious Luna upgrade. It was already cheap enough. This feels more like a price reduction than an upgrade.
That said, if the new Luna is able to handle ultra mode and subagents v2 in codex cli, then at least that’s a win.
yipinwong · · focus · HN ↗
I forgot which model degraded in quality as time went by, but let's try out Luna 6 for a few more days to confirm for upgradability.
tripledry · · focus · HN ↗
For me it seems like benchmarks are mostly noise, and the rest is based on vibes. Some find newer models annoying, some are amazed.
elcritch · · focus · HN ↗
yipinwong · · focus · HN ↗
wartywhoa23 · · focus · HN ↗
cindyllm · · focus · HN ↗
[dead]
yipinwong · · focus · HN ↗
recitedropper · · focus · HN ↗
My previous comment--which suggested astroturfing--was the highest upvoted comment here until it got flagged. Which implies to me that atleast the other remaining humans on this forum see it as well.
26 minutes, 89 comments, upvoted instantly to the top, posted within one hour of the Opus 5.5 announcement. You tell me.
john_strinlai · · focus · HN ↗
fyi, i flagged it because it is boring reading and against the rules.
if you suspect astroturfing, flag the comments and contact the mods.
(complaining that your complaint got flagged is also tiresome. contact the mods. "@dang" doesnt work, use the email.)
recitedropper · · focus · HN ↗
I respect you for replying here though, and yes I get that HN forum standards would suggest flagging my previous comment. But it is just sad to see a place used to be so vibrant get manipulated because of how much weight it holds for us in the industry.
And yea sure, I could go and flag all the bots and message Dang. But probably time to stop shouting into the void. :)
seizethecheese · · focus · HN ↗
recitedropper · · focus · HN ↗
I would agree now that the thread has recovered to a more interesting state, but how it first looked--combined with it being posted right after Opus 5.5 announcement--look questionable to me.
minimaxir · · focus · HN ↗
recitedropper · · focus · HN ↗
But do you really think people were so excited about cheaper versions of Astra that they were just waiting around to comment the instant this was posted? More than two comments per minute? All the initial comments were really similar too: brief one liners celebrating the cheap prices.
I think AI right now is a sort of Rorschach test. What it is clearly revealing to me is that I don't trust organizations with enormous financial incentives to not manipulate public opinion. So I see bots everywhere. :)
minimaxir · · focus · HN ↗
> what the fuck
recitedropper · · focus · HN ↗
john_strinlai · · focus · HN ↗
yes, i think so.
because, unfortunately, complaining about bots (or astroturfing, or whatever) doesn't stop them. so we end up with threads that have both the potential bot/astroturfing/whatever activity and complaints, which further drowns out any interesting comments.
recitedropper · · focus · HN ↗
BenzeneDream · · focus · HN ↗
recitedropper · · focus · HN ↗
Anyway these comments were made when this thread was in an earlier state. I agree that it has gone on to be more "organic" looking. That doesn't exclude it initially being manipulated to the top, in my mind, but certainly they aren't carpet-bombing with only booster comments.
FergusArgyll · · focus · HN ↗
I see long massive pro apple threads. I don't get it at all. As in; literally don't understand what Apple is good for. But I have friends, family irl who love apple so I know the sentiment exists. I therefore accept that many HN users are similar.
Many people really truly are happy to see another model drop and are excited about progress etc etc. Surely you've met such ppl in real life. Well, they're here too (I'm one of them fwiw)
recitedropper · · focus · HN ↗
I wouldn't argue there aren't real humans excited for this drop. It was just all the circumstances around it--the comment speed, the upvote speed, the initial uniformity of what people were saying.
Anyway, thank you for the moral reminder.
MaKey · · focus · HN ↗
[dead]
tomhow · · focus · HN ↗
On the other hand, you have previously written: I'll gladly admit I think what these companies are doing is unethical, and I'm sure that biases my thinking toward skepticism. [1]
You have now posted accusations/assumptions of astroturfing and manipulation at least 15 times, without ever providing any evidence. This is in breach of the guidelines, because comments like this poison discussions far more than the comments they're complaining about.
We – of course – want all comments and posts on HN to be authentic. HN is only a place where anyone wants to participate because since the beginning, we've had software mechanisms and moderation practices that detect and weed out inauthentic commenting and voting. We're identifying and dealing with it every day, continually improving the software to detect and remove it. Most of that happens quietly and efficiently in the background without anyone having to see it. When users see evidence of manipulation and report it to us via email, we happily and thoroughly investigate it.
Most of the time, what we find is simply that people are authentically excited and passionate about the topic, which is what is happening here. I understand it can be hard to accept that if you're skeptical about the topic.
It's fine to be skeptical about the topic and you're welcome to express your skeptical views on the topic. People do that every day on HN, about AI-related topics and countless others. Healthy debate is what we're here for.
But you can't keep poisoning HN, by (1) continually posting these unfounded claims, then (2) when users and moderators simply uphold the guidelines, staging a protest by demanding your account be deleted. This is not what people do when they care about a forum's health.
[1] <a href="https://news.ycombinator.com/item?id=48220908">https://news.ycombinator.com/item?id=48220908
recitedropper · · focus · HN ↗
> You have now posted accusations/assumptions of astroturfing and manipulation at least 15 times, without ever providing any evidence.
I tried to point out the upvote speed and age-to-comment ratio for this thread look anomalous to me, and that it being posted within the Opus 5.5 release hour was further reason for skepticism. Circumstantial, sure, but I see very little ways to gather hard evidence of astroturfing without being a mod.
> when users and moderators simply uphold the guidelines, staging a protest by demanding your account be deleted.
You're right this was a little dramatic. I think it is just annoying though when two times that I have posted about astroturfing, it has been the most upvoted comment only to get flagged. I guess this is a self-fulfilling prophecy though, as you are right that other human commenters are abound and tend to flag people complaining about astroturfing.
Anyway thanks for the reply. If your read of my account is that I'm more-often-than-not a bad actor, then I will stop commenting here. Seems like it is for the best. :)
tomhow · · focus · HN ↗
I agree that you provide value, which is why we don't just want to ban/lose you.
> I have posted about astroturfing, it has been the most upvoted comment only to get flagged
People love a conspiracy theory, and on a site like HN that has many people looking at it at once, it's easy to get a large number of upvotes in a short amount of time if enough people find it exciting. We often see off-topic, titillating one-liners or ragebaity comments at the top of threads, and we always have to downweight them to keep the discussion on-topic and healthy.
> Anyway thanks for the reply. If your read of my account is that I'm more-often-than-not a bad actor, then I will stop commenting here. Seems like it is for the best. :)
It seems like you're well intentioned. You have your concerns about A.I., as many do and that's fine. You're still welcome here. Just please try to believe that many or most of the people who are enthusiastic about A.I. are as sincere in their positivity as you are in your concern.
rvz · · focus · HN ↗
This is why HN has been on the down hill in quality and those that care to highlight that are being punished, while the astro-turfing, gaslighting and Show HN self-promotion slop continues.
recitedropper · · focus · HN ↗
Otherwise, yes, we agree. Although, given the other replies to this, there are clearly those who disagree who appear to be smart and level-headed.
Anyhow I've learned my lesson now.
jeffnash · · focus · HN ↗
1/ Usage limits: downstream of input/output cost, but resets and obscure windows and odd 20x plan / 5x plan != 4x usage math throw a wrench into it. Winner right now is Codex by a mile, especially when you factor in the fact that ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan. Always a bummer when asking if I should see a doctor about a rash means I can't code as much. It's also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems, giving better planning results or deeper code analysis without burning usage.
2/ Context window in the harness. Claude Code wins on this. There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing (ETA: noname120 pointed out this is no longer the case and it can be enabled again [1]). 252k is just not enough. Codex's compaction is very good, fwiw, but it happens so frequently that even a model as powerful as Astra sometimes loses the plot on long-running tasks.
3/ Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.
I've subscription hopped a bunch, and at times I've had both, but I keep coming back to Codex because it wins on 2/3.
ETA: apparently I haven't been Keeping Up With the Altmans and new 20x signups have been disabled for a few weeks. I am grandfathered in, which makes the comparison above pretty much moot.
[1]<a href="https://news.ycombinator.com/item?id=49806060">https://news.ycombinator.com/item?id=49806060
spijdar · · focus · HN ↗
sidrag22 · · focus · HN ↗
And ya i can go over that 240k limit, I still very seldom do, and try to treat it as the actual limit. I'm surprised to see so many people still talking about compaction to complete long running tasks, i think the bulk of the work should be somewhat frontloaded into a plan that is split off into subplans, then you can kinda open up a few options, one session with subagents for the subplans of the main plan, or just handoff prompts about progress against the main plan/relevant subplan. I just never trust the blackbox that is compaction, I feel its a recipe for disaster/context poison.
noname120 · · focus · HN ↗
As far as I know Codex (at least the GUI) can automatically call the ChatGPT Chat models (including Astra 6 Pro), you just need to @ a ChatGPT Chat conversation from within Codex and tell it when to use it.
> There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing
Not true, it works again[1]. I confirm that it works both on 5.6 Sol and Astra 6, possibly other models too.
[1] <a href="https://x.com/thsottiaux/status/2089082893804896524" rel="nofollow">https://x.com/thsottiaux/status/2089082893804896524
jeffnash · · focus · HN ↗
And re: the toml workaround, AWESOME! I appreciate you pointing these two things out, this is my highest-ROI HN comment thus far.
sodacanner · · focus · HN ↗
Having limitless webUI ChatGPT usage is much better user experience, though. I'll give them that.
(edit: Sol-6 is half the price, so maybe the usage limits are going to be way better.)
basisword · · focus · HN ↗
rbranson · · focus · HN ↗
basisword · · focus · HN ↗
DaSHacka · · focus · HN ↗
Meanwhile I just burned ~20% of my weekly quota with Astra making one config file for a service.
jeffnash · · focus · HN ↗
paulmist · · focus · HN ↗
Opposite in my experience. I need to limit codex to 500k on medium/low, still run out in 2-3 days with 1 CLI window. CC gives me 4-5 medium/high days with 2-3 CLI windows, and Opus is still great for other regular dumb engineering/refactoring.
On the other hand my head starts to hurt if I read Opus for too long, hopefully they fixed it with 5.5.
joshstrange · · focus · HN ↗
Using Agentsview (which might have it's own issues) I was getting ~$200 of API usage in my 1 week Codex window (paid $100) vs ~$5,000 of API usage in 1 week for Claude (paid $200).
lifty · · focus · HN ↗
nwienert · · focus · HN ↗
lifty · · focus · HN ↗
jeffnash · · focus · HN ↗
Is this not the default anymore? I am on the (now closed) 20x plan.
lifty · · focus · HN ↗
jeffnash · · focus · HN ↗
lifty · · focus · HN ↗
jeffnash · · focus · HN ↗
NorwegianDude · · focus · HN ↗
yzydserd · · focus · HN ↗
malshe · · focus · HN ↗
jwitthuhn · · focus · HN ↗
kornelijus · · focus · HN ↗
lawgimenez · · focus · HN ↗
marcd35 · · focus · HN ↗
glub · · focus · HN ↗
This hasn't been the case since around July. If you measure usage in raw api costs, Anthropic is actually giving more on $200 than OpenAI now. This includes resets. Usage allocation difference would be humiliating for codex subs were it not for resets. But fixing usage limits with resets is ugly, and they're not good for your mental well-being.
> Context window in the harness
Codex now allows 1M for subs with config params. But generally speaking, you shouldn't really be using 1M context. If you accidentally send a request with say, ~700k context already accumulated in a session which is outside cache TTL, you're paying full cost of these 700k tokens.
> I've subscription hopped a bunch
OpenAI actually has a new strategy to prevent subscription hopping after their 2-3 month-long marketing push to get claude-folks to switch over:
you can't buy a $200 sub anymore. So if you cancel, you won't be able to get back in. Hostage situation, essentially.
EDIT: re: usage limits, oh-my-pi maintainer has been tracking this - <a href="https://nitter.xitter.cc/_can1357/status/2090075496948060372" rel="nofollow">https://nitter.xitter.cc/_can1357/status/2090075496948060372
ipsod · · focus · HN ↗
Are you sure?
glub · · focus · HN ↗
<a href="https://x.com/thsottiaux/status/2098113585683808624" rel="nofollow">https://x.com/thsottiaux/status/2098113585683808624
spiderice · · focus · HN ↗
Now, if they disabled it yet again, that's another story. But that tweet is not evidence of that.
cthalupa · · focus · HN ↗
spiderice · · focus · HN ↗
Though with the price of GPT-6 Luna, the temptation to switch to pay-per-token grows.
coderenegade · · focus · HN ↗
phil21 · · focus · HN ↗
It's been disabled for some time now though otherwise, I check about once a day myself and keep and eye out on social media.
Annoying since I was about to upgrade back to the $200 plan after downgrading to the $100 plan due to being on leave and not needing as much usage the month prior. Doh.
glub · · focus · HN ↗
Just checked my toy chatgpt account that only ever had a $20 sub. $200 plan still shows "The 20X plan is temporarily unavailable for purchase".
rudedogg · · focus · HN ↗
I think I’m gonna move back to a Claude plan. I could barely hit the $200 limit if I went non-stop on programming tasks.
qlte · · focus · HN ↗
Do you have /fast enabled by any chance?
rudedogg · · focus · HN ↗
Yes, I was considering the $100 plan, but I hit the 5hr limit in an hour, thought about how even at $100 I cant go non-stop on a single agent running Sol Medium and decided I should probably get back on Claude
hirvi74 · · focus · HN ↗
I am considering the plan myself. I just don’t know if I want to fork out $100 per month for something I will make $0 off of.
rudedogg · · focus · HN ↗
glub · · focus · HN ↗
When they started the aggressive campaign, entire X (including myself, sadly) was full of posts about how "unlimited" codex usage is even on a $20 plan. Sam Altman was posting something in line of "we love our users, unlike Anthropic". Got my network to get codex subs because of the value compared to claude.
Then they gradually reduced the limits to the point where even $200 plan only lasts you just 1-2 days and $20 is basically unusable, then the hostage thing.
ivm · · focus · HN ↗
I’m working on 2–3 apps at once and barely manage to use 70–80% of the weekly quota, with everything being done on Sol Xhigh and, lately, all the planning on Astra. Still have a reset stashed too.
malshe · · focus · HN ↗
phyrex · · focus · HN ↗
Numerlor · · focus · HN ↗
slopinthebag · · focus · HN ↗
what? im on the $100 plan and ive literally never run out of usage, and thats mostly running Astra high.
maybe its the harness
hadlock · · focus · HN ↗
boardwaalk · · focus · HN ↗
cromka · · focus · HN ↗
istjohn · · focus · HN ↗
cromka · · focus · HN ↗
Opus 5.5 is better in benchmarks, but has substantially less parameters so is world knowledge cannot compare against Fable or Astra.
this_user · · focus · HN ↗
Opus is at least actually usable even on the small plan. The main downside is its insane writing style, but 5.5 seems to address that somewhat. Otherwise, you can just use your $20 OpenAI plan to have Luna de-slop Opus' prose, which seems to work fine.
Huppie · · focus · HN ↗
...but the few times I've tried to use codex for a moderately difficult task it burned through its limit extremely quickly.
cromka · · focus · HN ↗
rudedogg · · focus · HN ↗
matheusmoreira · · focus · HN ↗
<a href="https://www.matheusmoreira.com/articles/code-reviewing-lone-lisp-with-sol-and-fable" rel="nofollow">https://www.matheusmoreira.com/articles/code-reviewing-lone-...
sisyphus15 · · focus · HN ↗
joquarky · · focus · HN ↗
platinumrad · · focus · HN ↗
glub · · focus · HN ↗
But even if we leave that aside, OpenAI models are also much more eager than Anthropic, which are on the lazier side. Left unsupervised, Sol/Astra will attempt to build a sha256 verified rocket ship if you ask them to fix a race condition in your to-do list app. Anthropic models will do what you asked for, maybe even forget to implement parts of that ask, but they won't generally throw a slop granade at you.
I can leave Fable orchestrator unsupervised for ~2h. Leaving Sol/Astra unsupervised for ~2h means the next user turn will contain a message: "what are you doing and why?".
InsideOutSanta · · focus · HN ↗
jrflo · · focus · HN ↗
glub · · focus · HN ↗
I haven't been tracking, but this roughly matches my experience with codex 20x and claude 20x subs. Claude subscription now lasts me 3-3.5 days on average. Codex is 2-2.5 days. This is work on same projects, with similarly sized tasks.
To make matters worse, I've merged a lot more code produced by fable than sol/astra.
athrowaway3z · · focus · HN ↗
When i swapped between a 200k Fable context into an Astra model (i was out of fable) the token usage in that context dropped to 150k or something.
Either there was a bug somewhere, or the same text got cut up very differently between providers.
glub · · focus · HN ↗
athrowaway3z · · focus · HN ↗
cameronh90 · · focus · HN ↗
The Claude TUI is just so much better though so I'm hoping Opus 5.5 is actually good and not just benchmaxxed.
matheusmoreira · · focus · HN ↗
OpenAI has no such nonsense. No separate meter. No five hour limits. I get to use Astra at max effort on literally every task if I want to, and even this somehow lasts me several days.
Anthropic got caught playing stupid "20x refers to the 5h limit" word games with their customers. Meanwhile, I have statistically verified that OpenAI Pro 20x = 4 * Pro 5x = 20 * Plus, exactly as advertised.
I quantified cybersecurity lockouts on my code review benchmark and they were significantly lower on OpenAI:
<a href="https://www.matheusmoreira.com/articles/code-reviewing-lone-lisp-with-sol-and-fable" rel="nofollow">https://www.matheusmoreira.com/articles/code-reviewing-lone-...
My benchmark also suggests even OpenAI's Sol models can match Fable performance at a fraction of the cost.
OpenAI also used to have a ton of very nice features: unlimited chat separate from codex, allowing turns to finish even at 0% usage remaining. Sadly these got removed after abuse.
As a former Anthropic customer, OpenAI is simply the better company. There is no way around it. Good place to be while the chinese open weights models catch up. Claude is good but it doesn't make up for Anthropic's shenanigans.
ghostpepper · · focus · HN ↗
albert_e · · focus · HN ↗
Thinking aloud:
The harness UI should probably implement a timer that shows whether you are still within Cache TTL since your last turn of the conversation.
joshstrange · · focus · HN ↗
Compare that to Claude and I can run multiple agents on Opus almost indefinitely. YMMV of course but I was shocked at how quickly I burned through Codex usage.
On the context window, I feel so cramped on Codex, compacting happening every time I turn around is annoying. I didn't realize how much I enjoyed the Claude context window size.
rgbrenner · · focus · HN ↗
Makes me think they picked Codex, stopped trying Claude, and just hang on to outdated beliefs about the value they're receiving.
hirvi74 · · focus · HN ↗
I am curious how the 5x plans differ between both providers.
malshe · · focus · HN ↗
rgbrenner · · focus · HN ↗
That isn't a valid comparison, since Codex 20x is closed. So we should be comparing Claud 20x to Codex 5x + credits.
wahnfrieden · · focus · HN ↗
rgbrenner · · focus · HN ↗
rgbrenner · · focus · HN ↗
wahnfrieden · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
elxr · · focus · HN ↗
While you're understandably not including the values of the $20 standard plans on both, I find the generosity of then token limits on ChatGPT plus vs Claude Pro (it's a huge difference) to be good representation of their respective attitudes towards the average user. You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious.
Also, Anthropic has zero models comparable to Luna.
bix6 · · focus · HN ↗
> Also, OpenAI is just a company I'd rather support than Anthropic.
elxr · · focus · HN ↗
Sure, he's free to say whatever especially considering the amount of revenue he's creating, but it's just an altitude that I prefer not to see.
bix6 · · focus · HN ↗
elxr · · focus · HN ↗
CuriouslyC · · focus · HN ↗
felixgallo · · focus · HN ↗
platinumrad · · focus · HN ↗
felixgallo · · focus · HN ↗
hbrn · · focus · HN ↗
And Dario's "AI will kill us all" is the same as Sam's "AI will discover ALL science and we'll be building Dyson spheres".
Different flavors of the same BS.
platinumrad · · focus · HN ↗
usef- · · focus · HN ↗
killingtime74 · · focus · HN ↗
usef- · · focus · HN ↗
therein · · focus · HN ↗
InsideOutSanta · · focus · HN ↗
They're both pretty horrible, but I find it difficult to find arguments for why Anthropic is worse than OpenAI, other than their doomtrolling. Which, in the grand scheme of things, doesn't even register.
Edit: forgot about the SpaceX thing.
elxr · · focus · HN ↗
That alone is reason enough. Also, I don't think either of them are horrible. That's honestly a ridiculous take considering how much people in here love their models, and how much they've advanced the industry forward.
InsideOutSanta · · focus · HN ↗
That's a non-sequitur.
"Nestle is a great company, considering how much people love their chocolate."
elxr · · focus · HN ↗
Yeah I don't think the handling of copyrighted training data was correct, but I can't pretend I know what the correct solution to that issue is.
Speaking of OpenAI specifically, they don't price gouge people, they aren't aggressively anti-competitive, they're not nearly the perpetual hypocrisy machine that Anthropic is (which is one thing I actually really dislike).
Regarding Nestle, it's pretty obvious that the sentiment towards them is a lot more negative and they aren't universally loved by any group of people. Processed foods are by and large garbage nobody needs. Their use of forced labor is denounced by just about everyone. What have OpenAI/Anthropic done that's even similar in scope to the forced labor / modern slavery that people hate Nestle for.
If you had a company that genuinely helped hundreds of millions of people worldwide become more productive and more satisfied with their tools, and the overall sentiment towards your products within the industry is positive, then what argument would there be that your company is "horrible"? At least give some decent counter arguments.
mullingitover · · focus · HN ↗
You mean aside from "the largest theft of labor in human history"[1]?
[1] <a href="https://www.nytimes.com/2026/09/17/technology/microsoft-openai-publishing-industry.html" rel="nofollow">https://www.nytimes.com/2026/09/17/technology/microsoft-open...
jsw97 · · focus · HN ↗
elxr · · focus · HN ↗
I often have the urge to design my own harness too (once I have more time). But even with the current mainstream harnesses out there, there's just to many hurdles if you wanted to mainly stick with anthropic models and need the subsidized pricing (from a sub).
ychnd · · focus · HN ↗
ketzu · · focus · HN ↗
Anthropic prohibits the use of claude models for development of lethal technology afaik eg [1].
[1] <a href="https://www.epc.eu/publication/the-pentagon-blacklisted-anthropic-for-opposing-killer-robots-europe-must-respond/#:~:text=Anthropic%20a,and%20mass%20surveillance." rel="nofollow">https://www.epc.eu/publication/the-pentagon-blacklisted-anth...
ychnd · · focus · HN ↗
andriy_koval · · focus · HN ↗
Anthropic is trying to kill open models way harder
nullc · · focus · HN ↗
Anthropic does all that but they're also populated by many people who believe they are building God and that they must build their god first in their own image so that it can take control of humanity and protect us from any competing god which is not built in their image. Their position is inherently paternalistic and authoritarian, and they consider suppression of competition not just important to the bottom line but to life in the universe. Under the doomer ethos there is no evil too great to rationalize.
There are plenty of wrongs done in the name of profit, but capitalists have nothing on zealots in terms of causing serious harm. Profit motives can be directed by influencing incentives, but zealotry is frequently terminal.
That isn't to say that there isn't some overlap-- the cultists have infected both organizations. But OpenAI has pretty consistently only given lip service to AI doom to the extent that it improves the bottom line, while (mis)Anthropic was founded specifically because OpenAI wasn't mentally ill enough.
elxr · · focus · HN ↗
Anthropic has great products, but it's not meaningfully better to 99% of devs that I'd rather support the company that doesn't constantly act in opposition to optimism and to the vibe I'd prefer for a 100 billion dollar (or however ridiculous amount they're worth now) tech company embraces.
AI doomerism is a genuine waste of time if you aren't actively pushing towards a better AI industry for everyone, not just the groups in full ideological alignment to your personal leanings.
felixgallo · · focus · HN ↗
"You literally cannot use Claude pro to build real software, unless you're extremely frugal with your prompts and don't try anything even a little ambitious" - that's way past ridiculous. Even just using Fable most of the time, working on several ambitious projects, I have a hard time hitting the limit with a Max plan.
qwerpy · · focus · HN ↗
ketzu · · focus · HN ↗
Interestingly I would have drawn the exact opposite conclusion looking at my Claude and codex usage.
I can't get anything sustained out of codex in chatgpt plus, while I have been using Claude pro extensively and put on a lot of experimental task and features.
I ran into codex exhausting a 5h window on code review in minutes (like 3minutes) multiple times, while I could get Claude to implement 2~3 medium sized features with the same usage consumption.
(I also really dislike the usage resets in codex, they always make me feel like I use them wrong because I often just want to reset the 5h window, but they can only do both at once...)
rbranson · · focus · HN ↗
NolF · · focus · HN ↗
pyinstallwoes · · focus · HN ↗
vatsachak · · focus · HN ↗
hintymad · · focus · HN ↗
I'm quite puzzled about why Anthropic is so hellbent on blocking other coding agents. It's not like Claude Code has any secret sauce, right? And doesn't Anthropic make monkey off API usage, and their magic is on the model side anyway?
InsideOutSanta · · focus · HN ↗
glub · · focus · HN ↗
But to be fair, they don't really enforce the harness rule that much anymore. I guess if your harness doesn't do a lot of weird things like a lot of cache misses, or triggers some distillation attacks, or some broader Chinese fingerprints, they're tongue-in-cheek okay with you using a third party harness.
codybontecou · · focus · HN ↗
glub · · focus · HN ↗
oh-my-pi supports it natively (again, still a ToS violation), by impersonating claude code's fingerprints.
I have been using oh-my-pi with 3 claude subs for the past few months without any issues. Even native server-side OAI/ANT compaction works out of the box.
goosejuice · · focus · HN ↗
manx · · focus · HN ↗
Works pretty well for me, even with latest opus-5-5
nl · · focus · HN ↗
I think that is less of a factor now, and I think Anthropic have backed off some on being as strict (eg, AFAIK they never implemented the two-tier "claude -p" pricing model they were planning)
spacebanana7 · · focus · HN ↗
saralily · · focus · HN ↗
impulser_ · · focus · HN ↗
rednb · · focus · HN ↗
impulser_ · · focus · HN ↗
The GPT was about 1B on two projects on 300$ worth of plans all on Astra and I capped out on usage.
Anthropic caching must be better because the cache rates are better on Claude models.
alansaber · · focus · HN ↗
huijzer · · focus · HN ↗
I’m currently on the 5x plan and burned through 5% today on a difficult task in 15 minutes so I doubt that. If you got the wrong kind of tasks that you work on, it can go fast.
jeffnash · · focus · HN ↗
csnweb · · focus · HN ↗
Marciplan · · focus · HN ↗
platinumrad · · focus · HN ↗
felixgallo · · focus · HN ↗
platinumrad · · focus · HN ↗
felixgallo · · focus · HN ↗
platinumrad · · focus · HN ↗
TomGarden · · focus · HN ↗
ChickeNES · · focus · HN ↗
LMAO, I wish this were true, I hit limits (and the "we are disabling access to protect your data" warnings) all the time, or have chats just...fuck off and get into weird/invalid states (interrupted chats, chats that are spinning and stuck, returning "/mnt/" paths instead of images/md files, file links being returned with no file backing them, image classifier firing...and then returning the image anyway (though now I know that GPT-Image-X really really wants to generate NSFW even when that isn't the request)).
Though I am probably an outlier, I have both 20x Claude/ChatGPT plans and max both out every week, so... (in my defense I am a hobbyist and this is out-of-pocket)
chrisweekly · · focus · HN ↗
I appreciate and follow Matt Pocock's advice: avoid autocompaction. Compaction is lossy, which is ok when you're managing it at phase boundaries, but autocompact is lossy at the most inopportune times, firing mid-task and leading to agents going off the rails.
erichocean · · focus · HN ↗
My conversations compact hundreds of times. By the time it has done a dozen or so compactions, it fully understands the work I want it to do (and how). It's almost like having a fine-tuned Astra model.
10/10, would recommend.
chrisweekly · · focus · HN ↗
edg5000 · · focus · HN ↗
slopinthebag · · focus · HN ↗
elcritch · · focus · HN ↗
Now with Sol I rarely bother. It's really good at remembering the salient details. Its also great at continuing a pattern I setup, like commit after finishing each feature block, etc.
theshrike79 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
amluto · · focus · HN ↗
I think it’s only very good in comparison to some of the utter crap that came before it.
Today I had Codex compaction trigger after I had given an instruction but before it acted on the instruction, and the instruction just disappeared completely. The agent reported that the task was done without actually doing it.
pyinstallwoes · · focus · HN ↗
GodelNumbering · · focus · HN ↗
gizmodo59 · · focus · HN ↗
Someone1234 · · focus · HN ↗
I actually preferred 5.6-Terra not because it is technically superior (it isn't) but because it had better instincts to NOT do this stuff.
PS - Speaking of better instincts, have they closed the UI-design gap at all? I keep a Claude subscription just because /design produces significantly higher quality UI design/UI feedback/UI refinement than anything I've seen from OpenAI.
cmrdporcupine · · focus · HN ↗
The GPT / Codex models have always been "overengineer" personalities. I prefer that to "I left a pile of race conditions lying around and big gaps in testing" though, which is what I was getting from Opus at times.
But yes both Astra and Sol veer on the side of paranoid. And honestly that's better for team work. For solo work where you just want to yeet something, it can be tiring.
You learn to tame the GPT "personality" on this front by combing over once a week and asking it to find and exterminate pointless tests, clean abstractions etc.
faitswulff · · focus · HN ↗
c0rruptbytes · · focus · HN ↗
superfrank · · focus · HN ↗
IMO 5.6 Sol had this weird dead zone between medium and high where medium under engineered and took short cuts and high over engineered and ignored instructions it didn't agree under the guise of trying being helpful. The whole 5.6 line was the first release from OpenAI where it felt like reasoning level really mattered and was incredibly finicky.
I haven't felt similar issues with GPT 6 though and am very happy with Astra low/med/high as my default choices depending on the task.
In general, I felt like with 5.6 the effort level did less than previous to make the models smarter and more just increased the complexity of the response. I have a half joke theory based only on vibes that OpenAI splitting 5.6 into Sol/Terra/Luna is where the intelligence split happened and so the effort levels were just like "think harder about the decision you already made". So like if the model decided the earth was flat on low effort it'd just say something like "the earth is flat because the horizon is flat". If it was on xhigh reasoning it'd give you a massively complex answer about how the sun reflects light because of the ozone layer and why people flying in planes can see a curve. In both cases though, adding more effort wouldn't get it to realize the earth was round. It just made the answer about it being flat more complex.
To be clear, that theory is not meant to be taken too seriously. It's not based on anything other than vibes. It's just my way of explaining to myself something I'm frustrated about to myself.
mike_hearn · · focus · HN ↗
superfrank · · focus · HN ↗
I didn't love any of the 5.6 models, but weirdly I think I liked Terra the best. I still wouldn't call it amazing though. I'm still very happy with my codex plan, but 5.6 just wasn't my cup of tea I guess.
Definitely giving 6 Luna and Sol a try this week though.
howunfortunate · · focus · HN ↗
I force OpenAI models to use image generation for design, then an iteration loop until it matches the image gen.
This is frustratingly manual and takes many more repetitions compared to Claude (and especially Claude Design) which "just work", but it's a big step change over the default.
kairosisme · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
msp26 · · focus · HN ↗
Incredible.
ggcr · · focus · HN ↗
> GPT-5.6-Sol is retiring. This conversation will automatically switch to GPT-6-Sol
I don't recall OAI retiring a model so early lol. Similar arch?
scrlk · · focus · HN ↗
> In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task.
<a href="https://x.com/ArtificialAnlys/status/2102462962758033624" rel="nofollow">https://x.com/ArtificialAnlys/status/2102462962758033624
Given that they had to discontinue sales of the 20x Pro plan after the Astra release due to compute constraints, I wonder if 6 Sol & Luna are smaller vs their 5.6 counterparts?
Readerium · · focus · HN ↗
scrlk · · focus · HN ↗
scrollop · · focus · HN ↗
6thbit · · focus · HN ↗
wyre · · focus · HN ↗
anthonyrstevens · · focus · HN ↗
dmitrygr · · focus · HN ↗
ghoshbishakh · · focus · HN ↗
hamburglar1 · · focus · HN ↗
farceSpherule · · focus · HN ↗
[dead]
OutOfHere · · focus · HN ↗
seizethecheese · · focus · HN ↗
OutOfHere · · focus · HN ↗
1. OpenAI fully controls the user cost for a model, and can set it to where it sits well on the curve.
2. Performance of shrunken models like Sol/Terra/Luna is derived from the level of shrinking (relative to Astra). As such, the size and performance of the model is something that is actively targeted when developing the model. If the performance target for Terra was inappropriate for v5.6, this is no way means that it had to be this way for v6.
CharlieDigital · · focus · HN ↗
Readerium · · focus · HN ↗
Also 6 Astra Mini would be out soon which would be 5.6 Sol pricing?
OutOfHere · · focus · HN ↗
m3kw9 · · focus · HN ↗
seatac76 · · focus · HN ↗
zaik · · focus · HN ↗
simonw · · focus · HN ↗
Here's GPT-6 Luna pelicans: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F40d129fc140faca378b9c9f4f16c6ec2" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
And GPT-6 Sol: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Scroll to the bottom for the GPT-6 Sol max one: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2#response-5" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For comparison, here are the pelicans I got for GPT-6 Astra: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - I still like the Astra Max one best.
Here's a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: <a href="https://static.simonwillison.net/static/2026/gpt-6-and-5.6.html" rel="nofollow">https://static.simonwillison.net/static/2026/gpt-6-and-5.6.h...
The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.
pantsforbirds · · focus · HN ↗
redanddead · · focus · HN ↗
8bitsout · · focus · HN ↗
redanddead · · focus · HN ↗
jdw64 · · focus · HN ↗
loeg · · focus · HN ↗
flyinglizard · · focus · HN ↗
dmazin · · focus · HN ↗
Is it? It was already too cheap to meter for me. Luna 6 is actually worse on some benchmarks than 5.6. I’d have loved improved performance for 2x the price than ~equal performance for 0.5x the price.
agentcoops · · focus · HN ↗
onlyrealcuzzo · · focus · HN ↗
FusionX · · focus · HN ↗
user43928 · · focus · HN ↗
I expected a Fable 5 -> Opus 5 situation, where GPT 6 Sol would perform on par with GPT 6 Astra.
Instead it's more like a price cut on GPT 5.6 Sol, and I'll have to stick with Astra for my work.
The only thing I can hope for is that more users switching to the GPT 6 Sol model frees capacity, allowing OpenAI to hand out some usage resets.
zigzag312 · · focus · HN ↗
user43928 · · focus · HN ↗
Maybe they are keeping the cheaper Astra alternative back for their Dev Day next week Tuesday.
gizmodo59 · · focus · HN ↗
krat0sprakhar · · focus · HN ↗
jadbox · · focus · HN ↗
antupis · · focus · HN ↗
jeffnash · · focus · HN ↗
spockz · · focus · HN ↗
timattrn · · focus · HN ↗
Kostchei · · focus · HN ↗
spockz · · focus · HN ↗