GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
hlynurd · · focus · HN ↗
tedsanders · · focus · HN ↗
thejazzman · · focus · HN ↗
<a href="https://amphetamem.es/meme?id=the-simpsons_06_12_71&text=We%27re+allowed+to+have+one.&j=" rel="nofollow">https://amphetamem.es/meme?id=the-simpsons_06_12_71&text=We%...
afroboy · · focus · HN ↗
pkulak · · focus · HN ↗
This is a decent win though, if it really is better. 6-sol was really no good, at least in my work.
cmrdporcupine · · focus · HN ↗
Will see if this remedies things.
pkulak · · focus · HN ↗
prodigycorp · · focus · HN ↗
These moves all make sense when you take into account the enterprise market.
<a href="https://news.ycombinator.com/item?id=49889873">https://news.ycombinator.com/item?id=49889873
aaronbrethorst · · focus · HN ↗
t-sauer · · focus · HN ↗
nsingh2 · · focus · HN ↗
SirMaster · · focus · HN ↗
algoth1 · · focus · HN ↗
oh_no · · focus · HN ↗
RugnirViking · · focus · HN ↗
gradus_ad · · focus · HN ↗
nojito · · focus · HN ↗
I remember when bandwidth was super expensive and now it’s dirt cheap.
vanviegen · · focus · HN ↗
iAMkenough · · focus · HN ↗
Consumers are now saying the new pricing is not so great. <a href="https://news.ycombinator.com/item?id=49897236">https://news.ycombinator.com/item?id=49897236
djfjkfkffkkf · · focus · HN ↗
Razengan · · focus · HN ↗
It's not even anything controversial..
simlevesque · · focus · HN ↗
necovek · · focus · HN ↗
Razengan · · focus · HN ↗
neta1337 · · focus · HN ↗
alch- · · focus · HN ↗
bogrollben · · focus · HN ↗
Razengan · · focus · HN ↗
coderenegade · · focus · HN ↗
bendtb · · focus · HN ↗
2yrrr · · focus · HN ↗
It’s not happening and the imminent bust is coming. Strap in while the music gets turned up (dots, ipo etc) and people decide to leave the partayyy!
CamperBob2 · · focus · HN ↗
zigzag312 · · focus · HN ↗
tngranados · · focus · HN ↗
zigzag312 · · focus · HN ↗
First, the subsidies to consumers for electric vehicles in Germany apply to all cars, not just those built in Europe. This effectively subsidizes the competition from China.
As I didn't quickly find any source making the direct comparison, I asked LLM to research it (I apologize for for this, but doing it manually would take too much time).
EU Commission's investigation calculated countervailable subsidy rates for Chinese BEVs: see "3.10.3. Calculation of subsidy rates" for a aggregate subsidy rates.<a href="https://eur-lex.europa.eu/legal-content/EN/TXT/?qid=1738249752833&uri=CELEX%3A32024R2754" rel="nofollow">https://eur-lex.europa.eu/legal-content/EN/TXT/?qid=17382497...
CamperBob2 · · focus · HN ↗
Meanwhile, in the US, companies like Boeing, GM, and Intel will never be allowed to experience more than minor financial inconvenience before the government bails them out with protectionism, loans, and outright subsidies.
I just don't see a material difference between how the Chinese government treats their strategically-important industries and the way we do here in the West.
zigzag312 · · focus · HN ↗
Yes, but those are two separate things. Lower Saxony owning part of VW doesn’t mean VW gets extra public money because of it.
High subsidy rates alter market dynamics. Can we agree on that?
CamperBob2 · · focus · HN ↗
zigzag312 · · focus · HN ↗
Subsidies to manufacturers in Germany are much smaller than both: subsidies to consumers in Germany, and subsidies to manufacturers in China.
So, there are different subsidy types with different subsidy rates which all impact the market. Higher rates usually have a higher impact, but type of subsidy can also change what kind of an impact they have. Any comparison quickly becomes complex.
I posted some numbers in an other post [0], but the numbers are not exact.
Subsidy is also one of many factors. Of course you also need capable people and manufacturing capabilities which China also has. Subsidy alone is not enough, but if needed capabilities exist, subsidy at high rates can help tip the scales.
[0] <a href="https://news.ycombinator.com/item?id=49920502">https://news.ycombinator.com/item?id=49920502
ok123456 · · focus · HN ↗
mixdup · · focus · HN ↗
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
semiquaver · · focus · HN ↗
This is brain-rotted zitron-conspiracy territory, utterly at odds with reality.
ActionHank · · focus · HN ↗
We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.
CuriouslyC · · focus · HN ↗
ActionHank · · focus · HN ↗
We've gone from 80% in some places to 80% in some more places.
CuriouslyC · · focus · HN ↗
Any area that is verifiable will trend inexorably towards 100% over time. In unverifiable areas, it'll always be "80%" because the ubiquity of "AI" style erodes its value, and ">80%" for unverifiable things involves fashion, cachet and "vibes" that humans will probably never knowingly let it have.
ActionHank · · focus · HN ↗
coderenegade · · focus · HN ↗
usef- · · focus · HN ↗
Have you used recent ones on any large projects or coding issues? They've improved tremendously lately.
mixdup · · focus · HN ↗
helloplanets · · focus · HN ↗
With Opus 5.5, it doesn't seem like model improvement is plateuaing. And Fable 5.5 will likely be dropped this or next week.
You can only imagine what they've got going internally.
phoghed · · focus · HN ↗
arctic-true · · focus · HN ↗
famouswaffles · · focus · HN ↗
semiquaver · · focus · HN ↗
CamperBob2 · · focus · HN ↗
usef- · · focus · HN ↗
arctic-true · · focus · HN ↗
serf · · focus · HN ↗
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
mixdup · · focus · HN ↗
password54321 · · focus · HN ↗
delillos · · focus · HN ↗
password54321 · · focus · HN ↗
JacobAsmuth · · focus · HN ↗
LPisGood · · focus · HN ↗
theturtletalks · · focus · HN ↗
colechristensen · · focus · HN ↗
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
CuriouslyC · · focus · HN ↗
omalled · · focus · HN ↗
On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it.
[1] <a href="https://proceedings.neurips.cc/paper_files/paper/2025/hash/c7c2c3ab977316105f9c3b3fdf769f0e-Abstract-Datasets_and_Benchmarks_Track.html" rel="nofollow">https://proceedings.neurips.cc/paper_files/paper/2025/hash/c... [2] <a href="https://oeis.org/A000530" rel="nofollow">https://oeis.org/A000530
xienze · · focus · HN ↗
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
redanddead · · focus · HN ↗
azan_ · · focus · HN ↗
People were talking about plateau for years already.
sebzim4500 · · focus · HN ↗
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
djdjdkdkfk · · focus · HN ↗
RussianCow · · focus · HN ↗
robryan · · focus · HN ↗
epihelix · · focus · HN ↗
But coding-wise, models keep getting better and cheaper. You can train for code correctness in a way you can't train for legal correctness, and you can test your code in an agentic loop in a way you can't test a legal opinion.
Hence your alternative reality.
(All that said, 2023 was GPT-4 territory. GPT-4o wasn't released until 2024. No matter what question you're asking, I struggle to believe you wouldn't notice the difference between GPT-4 and the current frontier model set. You can download and run any number of sub-27B local models that will be better than GPT-4. The pace of change in this field really has been insane.)
rspeele · · focus · HN ↗
In law if a model makes a mistake it pretty much takes a human to catch it, which is a vastly slower, riskier, and more frustrating feedback loop. So it doesn't surprise me they can stay dumb in that field while getting massively more capable in coding with multiple "step change" releases in the past 365 days.
luma · · focus · HN ↗
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
dgellow · · focus · HN ↗
- it’s correct there isn’t much fresh data anymore
- it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive
- it’s correct the finances don’t make sense
But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)
moosehater · · focus · HN ↗
dumberquestions · · focus · HN ↗
[dead]
JacobAsmuth · · focus · HN ↗
If I have some ML workload to run I can buy $x of Blackwell chips or I can buy significantly less $ worth of Vera Rubin chips to get the same performance. That's the key thing to keep in mind when you're talking about financials.
agoodusername63 · · focus · HN ↗
never mind that theres no guarantee we'll get that mythical AI. Never mind that the societal reformations would also impact their revenue numbers.
dgellow · · focus · HN ↗
john_strinlai · · focus · HN ↗
do you think it will be exponential forever?
RobCat27 · · focus · HN ↗
spathi_fwiffo · · focus · HN ↗
Fabs.
Either needing more fabs, new types of fabs, retooling existing fabs.
All of that takes years.
maybe we can design our way out of that too. But, I suppose that would be the similar breakthrough you are mentioning.
[deleted] · · focus · HN ↗
[deleted]
sebzim4500 · · focus · HN ↗
I think it's fully possible that it continues being exponential for decades like Moore's law did (and still is depending on exactly what you measure)
trentnix · · focus · HN ↗
What a time to be alive.
neta1337 · · focus · HN ↗
trentnix · · focus · HN ↗
- plan youth soccer practices
- develop well-formatted soccer game substitution schedules
- build and ship software in languages I haven't used in 25 years on platforms I've never programmed for
- do meal planning and build shopping lists
- prepare grocery shopping carts
- solicit medical advice
- perform Garmin watch data analysis
- administer devices (with SSH access) using natural language
- avoid counterfeit soccer jersey purchases
- create "Warrior Cat" graphic novels
- make cartoon strips
- troubleshoot appliances
- manage finances
- review accounting ledgers
- diagnose malware infections
- so much more
And we do it all from a simple prompt that we can talk to if we choose.
I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming.
I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved.
FiberBundle · · focus · HN ↗
I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been.
senderista · · focus · HN ↗
semiquaver · · focus · HN ↗
robryan · · focus · HN ↗
digdugdirk · · focus · HN ↗
To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.
Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.
So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.
willchis · · focus · HN ↗
> "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_
famouswaffles · · focus · HN ↗
In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D asset creation are all things Astra is >> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong.
digdugdirk · · focus · HN ↗
And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
famouswaffles · · focus · HN ↗
Replacing white collar work would be worth dozens of trillions of dollars. Software is not the only valuable job that can be done on a computer.
OpenAI and Anthropic already have what it takes right now to become trillion dollar companies even if the above doesn't materialize.
Chatgpt is used by a billion people every week. Their ads program hit $1B ARR in 200 days. And Anthropic is growing so fast they're in track to hit an Annual Revenue Run Rate of $100B before they IPO.
>And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
How long until...you could say that about the capabilities of past models but OpenAI still dwarf everyone else in consumer usage, and Anthropic still have enterprise usage on lock. In the end, neither the billion+ users of gpt or the enterprise customers are going to give a shit. And specialized models often perform worse than generalized ones.
OliveronData · · focus · HN ↗
Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.
Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?
From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).
There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.
That could be cost cutting too, true, but really? That's the only explanation? And nothing else?
Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.
luma · · focus · HN ↗
These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.
neta1337 · · focus · HN ↗
dwaltrip · · focus · HN ↗
What distinction are you drawing?
OliveronData · · focus · HN ↗
RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that "small model feel" to it. The better smaller models get at coding the worse they get at everything else too.
Going from Sol 5.6 to Astra, Opus to Fable, you can still get that "larger model feeling," though less so. The bigger models can reference things that you would not have expected.
The distinction I'm making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL'd for, and that hopefully anything else doesn't get negatively affected. RL'ing for Javascript world for example did not improve the C world when working with the models.
dwaltrip · · focus · HN ↗
OliveronData · · focus · HN ↗
For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.
I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.
This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same/similar, and the model is just able to display its capabilities better.
Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model's local maximum, and bigger models are still smarter, even if they are not able to display it.
Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?
kqr · · focus · HN ↗
interestpiqued · · focus · HN ↗
dcchambers · · focus · HN ↗
[dead]
chamomeal · · focus · HN ↗
holbrad · · focus · HN ↗
nbardy · · focus · HN ↗
jorblumesea · · focus · HN ↗
it's also why there have been so many calls for regulation and slowdowns.
LeBit · · focus · HN ↗
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
nozzlegear · · focus · HN ↗
0cf8612b2e1e · · focus · HN ↗
Insane pricing pressure on the horizon.
jimbob45 · · focus · HN ↗
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
simianwords · · focus · HN ↗
minimaxir · · focus · HN ↗
eli · · focus · HN ↗
amelius · · focus · HN ↗
skulk · · focus · HN ↗
amelius · · focus · HN ↗
aleph_minus_one · · focus · HN ↗
amelius · · focus · HN ↗
mholm · · focus · HN ↗
condour75 · · focus · HN ↗
SkyBelow · · focus · HN ↗
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
Mkengin · · focus · HN ↗
phpnode · · focus · HN ↗
jesse_dot_id · · focus · HN ↗
anotha_one · · focus · HN ↗
[dead]
lxgr · · focus · HN ↗
dandellion · · focus · HN ↗
lxgr · · focus · HN ↗
sharpshadow · · focus · HN ↗
wg0 · · focus · HN ↗
Wheen · · focus · HN ↗
Edit: <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf" rel="nofollow">https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
ChromeUltron · · focus · HN ↗
LPisGood · · focus · HN ↗
mckirk · · focus · HN ↗
system2 · · focus · HN ↗
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
wg0 · · focus · HN ↗
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
copperx · · focus · HN ↗
andybak · · focus · HN ↗
thraway3837 · · focus · HN ↗
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?
senderista · · focus · HN ↗
white_dragon88 · · focus · HN ↗
[dead]
system2 · · focus · HN ↗
jonatron · · focus · HN ↗
anotha_one · · focus · HN ↗
[dead]
SwabbyNat74 · · focus · HN ↗
tjwebbnorfolk · · focus · HN ↗
colpabar · · focus · HN ↗
infamouscow · · focus · HN ↗
Aboutplants · · focus · HN ↗
scrollop · · focus · HN ↗
Luckily it's not a mistake as now we have access to . . . dots.
(and sol 6.1, it seems)
sockaddr · · focus · HN ↗
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
vividfrier · · focus · HN ↗
[dead]
geeky4qwerty · · focus · HN ↗
pythonaut_16 · · focus · HN ↗
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
killingtime74 · · focus · HN ↗
mattnewton · · focus · HN ↗
franzcoughka · · focus · HN ↗
[dead]
motoboi · · focus · HN ↗
mynameisjonny_ · · focus · HN ↗
agluszak · · focus · HN ↗
orbital-decay · · focus · HN ↗
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but it's probably too boring for some.
jchw · · focus · HN ↗
esafak · · focus · HN ↗
toasty228 · · focus · HN ↗
copperx · · focus · HN ↗
RGS1811 · · focus · HN ↗
copperx · · focus · HN ↗
RGS1811 · · focus · HN ↗
A common form of this failure is the model picking up on random wordings from earlier in the session (e.g. some comment it made to me in the middle of a response, that I never explicitly endorsed) and then treating these as hard commitments. Or over-interpreting a specific word choice or clumsy phrasing as if it were a "load-bearing" constraint on the task.
None of this clumsiness would be so problematic if the model didn't have such a strong drive toward autonomy. It's much like with people: there's no shame in not understanding what you're being asked to do, provided you ask clarifying questions. There's no shame in ignorance if it's wedded to curiosity. Benchmaxing has RLVRed curiosity and clarification straight out of these models. It sucks.
copperx · · focus · HN ↗
denysvitali · · focus · HN ↗
blmarket · · focus · HN ↗
az226 · · focus · HN ↗
MisterMunchkin · · focus · HN ↗
dannyw · · focus · HN ↗
ChromeUltron · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
Lapalux · · focus · HN ↗
fraywing · · focus · HN ↗
Astra is a pretty impressive model. Excited to try this.
gobdovan · · focus · HN ↗
Tadpole9181 · · focus · HN ↗
gobdovan · · focus · HN ↗
anotha_one · · focus · HN ↗
[dead]
Nevin1901 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
jeffybefffy519 · · focus · HN ↗
s3p · · focus · HN ↗
phist_mcgee · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
bopou · · focus · HN ↗
glimshe · · focus · HN ↗
SirMaster · · focus · HN ↗
scottyah · · focus · HN ↗
IshKebab · · focus · HN ↗
rs_rs_rs_rs_rs · · focus · HN ↗
paulryanrogers · · focus · HN ↗
How many DCs are devoted solely to gaming?
lp92 · · focus · HN ↗
HelloMcFly · · focus · HN ↗
rs_rs_rs_rs_rs · · focus · HN ↗
Yes but it adds up when you consider that just on Steam alone there are 200 million monthly active users.
[deleted] · · focus · HN ↗
[deleted]
paulryanrogers · · focus · HN ↗
empthought · · focus · HN ↗
paulryanrogers · · focus · HN ↗
Today physical discs are rapidly becoming a niche that newer consoles just won't have. PC games are almost never sold physically anymore.
rs_rs_rs_rs_rs · · focus · HN ↗
An entire planet. Just Steam alone has one or two hundres million monthly active users.
paulryanrogers · · focus · HN ↗
All of gaming is in the 300B USD range while just the CapEx of hyper scalars is already over 200B.
lbrito · · focus · HN ↗
rs_rs_rs_rs_rs · · focus · HN ↗
Yeah? Show me the big movements against computer gaming.
lbrito · · focus · HN ↗
The differences with AI are: 1) we are starting off (mid 2020s) from a baseline point of already being in a hopelessly shitty situation, past the 1.5C warming target; and 2) Electronics, chips, data centers etc were already a thing for a long time, but industry took _decades_ to ramp up production to pre-AI levels, and these things are used everywhere for a huge number of things. Now we're consuming electronics/data centers/water/power at an unheard-of rate, and for a single purpose (AI) with questionable benefits, besides the private interests of a handful of people.
JDups · · focus · HN ↗
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
IshKebab · · focus · HN ↗
Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.
rs_rs_rs_rs_rs · · focus · HN ↗
But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?
> I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Just Steam has 200 million monthly active users. Add Steam, PS, Xbox, and whole other devices having gpus and I'm pretty sure you at least 10x the energy consumption of all ai companies.
IshKebab · · focus · HN ↗
I dunno what you're not getting but a GPU to run games is like 200-500W. A GPU cluster to run Astra is probably more like 10kW.
Also gamers tend not to spin up dozens of other machines to also game for them.
otterley · · focus · HN ↗
lp92 · · focus · HN ↗
codehorses · · focus · HN ↗
IshKebab · · focus · HN ↗
Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.
Bolwin · · focus · HN ↗
sergiotapia · · focus · HN ↗
barrenko · · focus · HN ↗
Starlevel004 · · focus · HN ↗
dcchambers · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
Huge misstep releasing it.
algoth1 · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
tandr · · focus · HN ↗
iamdelirium · · focus · HN ↗
Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.
hyperpape · · focus · HN ↗
squidbeak · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
A_D_E_P_T · · focus · HN ↗
Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...
ekun · · focus · HN ↗
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
godwinson__4-8 · · focus · HN ↗
I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.
There are still rough edges of course. But try the MCP out and judge for yourself.
A_D_E_P_T · · focus · HN ↗
lukan · · focus · HN ↗
A_D_E_P_T · · focus · HN ↗
CuriouslyC · · focus · HN ↗
A_D_E_P_T · · focus · HN ↗
CuriouslyC · · focus · HN ↗
A number of others have done game/3d video benchmarks but this guy is probably the most prolific.
kroaton · · focus · HN ↗
2yrrr · · focus · HN ↗
therealdrag0 · · focus · HN ↗
minimaxir · · focus · HN ↗
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
TuxSH · · focus · HN ↗
bigwheels · · focus · HN ↗
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
Infinity315 · · focus · HN ↗
toasty228 · · focus · HN ↗
I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous
copperx · · focus · HN ↗
AndrewKemendo · · focus · HN ↗
squidbeak · · focus · HN ↗
edgyquant · · focus · HN ↗
AndrewKemendo · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
colinhb · · focus · HN ↗
rspeele · · focus · HN ↗
Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.
The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.
The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.
Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.
Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.
agar · · focus · HN ↗
this_user · · focus · HN ↗
rrvsh · · focus · HN ↗
chaostheory · · focus · HN ↗
rspeele · · focus · HN ↗
My biggest conclusion from this test was: the most efficient use of my weekly Astra budget is as a reviewer/consultant for work done by Opus. I don't have Astra write much code right now, but I do have it reading a lot of what Opus writes. Of course with the way the AI landscape shifts the balance could be the exact opposite 2 weeks from now.
Seeing how each model preferred its own flavor of code shows that, even from a "blind" fresh context, a same-model reviewer will still often look at the work of another incarnation of itself and go "yep that's how I woulda done it" and not be as likely to realize that there was an alternative path or implicit assumption/mistake in the work.
phoghed · · focus · HN ↗
They form these super strong opinions after a few prompts, then face reality over time.
People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.
beering · · focus · HN ↗
ex1fm3ta · · focus · HN ↗
mmis1000 · · focus · HN ↗
krzyk · · focus · HN ↗
Looks like 5.5 is the new 4.6
jauntywundrkind · · focus · HN ↗
dotancohen · · focus · HN ↗
peterbell_nyc · · focus · HN ↗
There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.
Starlevel004 · · focus · HN ↗
This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.
notatoad · · focus · HN ↗
I have the task to codex first, it took a couple back and forth prompts to define the project and then it worked for a bit and to took a couple more prompts before I decided it was good enough - not perfect, but close. It re-implemented some wrapper components in a simplified way that lost some of the UI, but it would work.
Opus 5.5 took the same prompt with no back and forth, it just went off a built a tool that takes pixel-perfect screenshots of exactly what my app looks like.
TuxSH · · focus · HN ↗
Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.
sobiolite · · focus · HN ↗
beering · · focus · HN ↗
rrvsh · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
dom96 · · focus · HN ↗
1 - <a href="https://bench.killswitch-lang.org" rel="nofollow">https://bench.killswitch-lang.org
zeroonetwothree · · focus · HN ↗
dom96 · · focus · HN ↗
Opus 5.5 fails the "understanding" tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers. Opus 5 gets it right.
Here are the outputs from both: <a href="https://gist.github.com/dom96/b5bce82b6e6c1ebd5271ed70ad941b49" rel="nofollow">https://gist.github.com/dom96/b5bce82b6e6c1ebd5271ed70ad941b....
Looking at that Opus 5.5 fails to deduce that the "hack statement" is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5's greater intelligence for what it's worth.
joshstrange · · focus · HN ↗
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
codewithcheese · · focus · HN ↗
redox99 · · focus · HN ↗
shimman · · focus · HN ↗
This is why these companies are struggling to make money, they're chastising their customers just like they've been chastising the human race.
trio8453 · · focus · HN ↗
It's very appropriate in the cases when you're holding it wrong. The fact that you're paying doesn't mean that you can't make mistakes or waste resources.
shimman · · focus · HN ↗
If this is how you want to get people on your side, I can understand why the entire country/human race are against these companies.
trio8453 · · focus · HN ↗
It's a product and if you're using it incorrectly, we can either
1. say so
2. pretend that you don't so to get/keep you on "our side"? or not say is because you're skeptical or hate it? (how does that last bit even follow logically?)
How is 2 better in any way for anyone involved?
crossroadsguy · · focus · HN ↗
Anonasty · · focus · HN ↗
jorblumesea · · focus · HN ↗
Aeolun · · focus · HN ↗
threecheese · · focus · HN ↗
I overused Astra in order to drain my weekly, figuring I'd have the reset. (not wastefully, I did get more work done)
Aeolun · · focus · HN ↗
seunosewa · · focus · HN ↗
nkmnz · · focus · HN ↗
edot · · focus · HN ↗
mattkenefick · · focus · HN ↗
I create a lot, but I can make a full month with Astra on the current Pro plan. What are you doing to spend that much?
redox99 · · focus · HN ↗
1 day is kind of generous, it probably lasts like 12 hours of running non stop.
apitman · · focus · HN ↗
antonvs · · focus · HN ↗
ChickeNES · · focus · HN ↗
Foobar8568 · · focus · HN ↗
Marha01 · · focus · HN ↗
antonvs · · focus · HN ↗
I've been using Gemini on development of a DNN training pipeline, and there's no way you can describe it as "dumb as hell". That description just makes it clear that you're not talking about the technical capabilities of the models, but about some sort of fanboy comparison from a parallel hype universe.
_davide_ · · focus · HN ↗
onlyrealcuzzo · · focus · HN ↗
No LLM will be cost effective if it's compacting this often. You have to find a way around it.
ngruhn · · focus · HN ↗
sally_glance · · focus · HN ↗
SyneRyder · · focus · HN ↗
Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:
<a href="https://x.com/thsottiaux/status/2089082893804896524" rel="nofollow">https://x.com/thsottiaux/status/2089082893804896524
There's at least a forum thread about it here:
<a href="https://community.openai.com/t/why-does-codex-report-a-258-400-token-context-window-for-gpt-5-6-sol/1394346/5" rel="nofollow">https://community.openai.com/t/why-does-codex-report-a-258-4...
gf000 · · focus · HN ↗
rrvsh · · focus · HN ↗
kaoD · · focus · HN ↗
bjord · · focus · HN ↗
yes, exactly
SyneRyder · · focus · HN ↗
These are often my best sessions - they're unattended overnight, because by then we have the specification figured out, and I can just leave Claude to build out the rest, making good choices if it does find gaps in the spec. I regularly go to sleep & wake up to an entirely new application completed. Claude never uses compacting in my sessions.
I haven't used GPT as much as I should have, so I'm prepared to be incorrect & out of date. It just intuitively feels like I wouldn't get the same from a 275K context window - maybe it uses lots of subagents? Even Deepseek & GLM have 1 Million context windows now, so it "feels" strange for people to actually prefer the 275K window. But that's just my intuition.
bjord · · focus · HN ↗
if you talk about them (in which you lean on an LLM as a sort-of independent employee) and conservative, chunk-based usage (in which you use the LLM as more of an extension of yourself), you're comparing apples to oranges
a predefined spec obviously reduces that gap but how much is highly dependent on the level of detail
jeremyjh · · focus · HN ↗
I’ve also found compaction not to be a problem when it does happen.
threecheese · · focus · HN ↗
jeremyjh · · focus · HN ↗
onlyrealcuzzo · · focus · HN ↗
rrvsh · · focus · HN ↗
Benjamin_Dobell · · focus · HN ↗
~/.codex/config.toml
AmazingTurtle · · focus · HN ↗
manmal · · focus · HN ↗
Gareth321 · · focus · HN ↗
It's crazy on Codex. I sometimes get just 2-3 turns before it compacts. It has forced me to use persistent project documentation for everything. Maybe that's not a bad thing but unless it reads all the documentation after every compaction (and uses half its cache), it goes off the rails. By comparison, Opus 5.5 is a breath of fresh air. It takes FAR longer to hit the cache limit and that means it keeps useful information in working memory far longer. I think this alone has resulted in a massive productivity and efficiency increase for me.
RugnirViking · · focus · HN ↗
exfalso · · focus · HN ↗
jmalicki · · focus · HN ↗
The longer your chat gets, the slower and more expensive it gets.
Subagents are expensive but they scale way closer to O(n) than O(n^2).
Have some agents make bug reports/feature requests/roadmaps (linear is very AI friendly), others coordinate, others work on grinding out an individual ticket.
If there is a good ticket-level description, it's a waste of time IMO to have a main agent do it, that should be an agent with fresh context that will do it better faster (the shorter the context, the better models are at using the context they're given).
jaktet · · focus · HN ↗
jmalicki · · focus · HN ↗
Whenever I see my main agent do a compaction, that to me is a clear sign I didn't have it delegate bounded tasks enough.
Still, I see no evidence Codex or Claude Code inherit full context of the main agent in subagents, I've always seen them be prompted, but this is something high priority on my list of unknowns to understand better...
KetoManx64 · · focus · HN ↗
Everyone else that uses memory files and a new conversation for each new sub project/feature rarely hit their weekly allotments.
verdverm · · focus · HN ↗
crazylogger · · focus · HN ↗
sscaryterry · · focus · HN ↗
[dead]
JimDabell · · focus · HN ↗
This is not even remotely true.
peterbell_nyc · · focus · HN ↗
If you're running 2-3 parallel agent session with a few sub agents and waiting for you to prompt them, you'll have a very different experience!
JimDabell · · focus · HN ↗
This is a tiny minority of people, not “most people”.
user43928 · · focus · HN ↗
With the 80% price cut, this is competitive with Opus 5.5 despite the subscription downgrade.
Additionally, it was said that existing 20x subscriptions retain the higher limits for some time.
I have seen you make these immature accusations that users here are OpenAI employees multiple times today.
sscaryterry · · focus · HN ↗
minimaxir · · focus · HN ↗
(Usage limits are entirely dependent on what you're doing with them. If you're not running it on 1 million LoC databases you can get a lot of mileage out of a 5x account)
vcryan · · focus · HN ↗
Also, a lot of this work is verification to ensure that AI generated code does what is intended and is safe to merge and deploy. That verification work is critical and uses a lot of tokens.
pvab3 · · focus · HN ↗
throwitaway222 · · focus · HN ↗
Guess not?
murbard2 · · focus · HN ↗
minimaxir · · focus · HN ↗
nimonian · · focus · HN ↗
anotha_one · · focus · HN ↗
[dead]
zf00002 · · focus · HN ↗
skybrian · · focus · HN ↗
alvis · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
slopinthebag · · focus · HN ↗
sehw · · focus · HN ↗
Aboutplants · · focus · HN ↗
Alifatisk · · focus · HN ↗
iosjunkie · · focus · HN ↗
mekpro · · focus · HN ↗
jdw64 · · focus · HN ↗
godwinson__4-8 · · focus · HN ↗
When is the alleged "safety" concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.
When do we get the next jump in capability? When is 6.1 Astra released?
ColonelPhantom · · focus · HN ↗
godwinson__4-8 · · focus · HN ↗
The coverage around 6.1 Astra seems deliberately playing into the dubious "safety" narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.
Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.
Without more details on the credibility of the safety concern this seems like a totally coherent action for customers to take. We shouldn't put up with teasing.
wren6991 · · focus · HN ↗
vb-8448 · · focus · HN ↗
tandr · · focus · HN ↗
az226 · · focus · HN ↗
TuxSH · · focus · HN ↗
objektif · · focus · HN ↗
vb-8448 · · focus · HN ↗
tultra · · focus · HN ↗
alexandroskyr · · focus · HN ↗
formvoltron · · focus · HN ↗
Alifatisk · · focus · HN ↗
gavin_gee · · focus · HN ↗
nicce · · focus · HN ↗
thefounder · · focus · HN ↗
the_duke · · focus · HN ↗
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
nxc18 · · focus · HN ↗
user43928 · · focus · HN ↗
The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.
I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.
nxc18 · · focus · HN ↗
I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.
holbrad · · focus · HN ↗
I haven't used it much yet, but I have much higher hopes for Sol 6.1, as it seems to be based off of a completely different base, it's not just a tune.
sebzim4500 · · focus · HN ↗
Astra is clearly far better than anything prior though, so I'm not sure what you mean really.
Hammershaft · · focus · HN ↗
Am I misinterpreting this, or did OpenAI clearly nerf GPT-6 Sol on the 23rd.
user43928 · · focus · HN ↗
The chart shows GPT-5.6 Sol and a surprisingly large drop in performance when the switched it over to GPT-6 Sol.
sigbottle · · focus · HN ↗
I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".
I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.
> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).
nxc18 · · focus · HN ↗
On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.
sigbottle · · focus · HN ↗
Yes, still running into this, but surprised about this
> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.
But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.
For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.
moshegramovsky · · focus · HN ↗
moshegramovsky · · focus · HN ↗
Here's a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That's a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn't. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn't follow existing API patterns.
jstummbillig · · focus · HN ↗
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
the_duke · · focus · HN ↗
phoghed · · focus · HN ↗
Codex itself seems to have a regression. You can see clearly the token use changing significantly coincides with a score drop
nicce · · focus · HN ↗
copperx · · focus · HN ↗
Marha01 · · focus · HN ↗
Eridrus · · focus · HN ↗
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
phoghed · · focus · HN ↗
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.
btbuildem · · focus · HN ↗
wkcheng · · focus · HN ↗
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
equinumerous · · focus · HN ↗
trentnix · · focus · HN ↗
chronogram · · focus · HN ↗
r0l1 · · focus · HN ↗
setnone · · focus · HN ↗
ozgung · · focus · HN ↗
bitexploder · · focus · HN ↗
NorthSouthNorth · · focus · HN ↗
sunaookami · · focus · HN ↗
stldev · · focus · HN ↗
For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.
For modeling and artwork, Astra has been great routinely outperforming Kimi.
This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.
I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.
rrvsh · · focus · HN ↗
I had to switch back to 5.6 Sol after trialling 6 Sol for like 3 days - I was getting insanely annoyed at how misaligned it is. Will try 6.1 but not very high hopes
keyle · · focus · HN ↗
I'll just leave this here: <a href="https://marginlab.ai/trackers/codex/" rel="nofollow">https://marginlab.ai/trackers/codex/
dannyw · · focus · HN ↗
And, is it really even an oligopoly anymore? Open weight models are incredibly competitive in every way; whether you want to use US providers, Chinese official providers, self host, etc.
moshegramovsky · · focus · HN ↗
I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30/40 minutes) to do simple things. As a result, Astra didn't get much done. It needs the same small implementation slices as GPT 5.5/others, but was much slower and didn't generate better results. (On a complex infra project/across a large codebase.)
It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn't seem to be better than 5.5 at most programming jobs.
I'm on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can't get coffee fast. Like I can't send an email fast.
OpenAI is making some excellent products for sure but I'm not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren't good at working autonomously on large codebase situations. Just because something compiles doesn't make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe/side app and then worked through the design there. Um, what? Not that it's invalid to do this but I actually have to test in the live codebase or I can't possibly say that something is working.
Just because you can, doesn't mean you should.
jrflo · · focus · HN ↗
skeptic_ai · · focus · HN ↗
soulofmischief · · focus · HN ↗
What was a pleasant and productive experience is becoming increasingly frustrating and draining.
beebmam · · focus · HN ↗
diego_sandoval · · focus · HN ↗
GPT 6 needs to be babysit, otherwise it starts doing ridiculous things.
4b11b4 · · focus · HN ↗
jsw97 · · focus · HN ↗
pampas · · focus · HN ↗
jeffybefffy519 · · focus · HN ↗
koyote · · focus · HN ↗
I've never seen such a large degradation in intelligence in a model until I tried out Sol 6 after having used 5.6 almost exclusively for several weeks.
twotwotwo · · focus · HN ↗
And, of course, GPT-6 came out as Anthropic fixed a bunch of stuff with their models -- faster (via fewer tokens, and TPS for Sonnet), easier to work with, better results, cheaper (via pricing and, again, fewer tokens). I don't know if the timing and the sudden improvement on Anthropic's side sharpened the vibes comparison this round, but Internet opinion went pretty clearly to Anthropic.
FrontierCode's results make it look like the lower two effort settings of Sol-6.1 may be better options than Sonnet or Opus on low, but you might be better off with Opus than with Sol's high settings.
One thing I don't think any of this reflects is that many well-specified coding tasks, including the self-testing and doing research and tracing out dependencies and so on, aren't really bleeding-edge now: Luna-5.6 and small open models handle them fine. Stuff like "why is this box dropping connections?" or "here's a thing I want you to model/figure out" can benefit from bigger models. But far from everything does!
laurels-marts · · focus · HN ↗
I tried out fable 5.1 the day it was released and coming from gpt-5.6-sol I was truly mind blown (both in terms of code and prose it was generating - outputs I could finally enjoy reading and looking at).
Then when opus 5.5 came out, again same thing + far cheaper and faster.
I went from using OAI exclusively the entire year to a point now where i haven’t touched one of their models in at least a few weeks now.
I think OAI has lost the plot. OAI models simplify have no taste. And I don’t mean in front-end design way (although that too). They have no taste in how the model writes code, how it writes prose, how it writes in-line comments, how it writes documentation, or how it even picks variable names. There’s just no taste throughout.
Anthropic models are very thoughtful and have so much taste all around.
stasomatic · · focus · HN ↗
jp_gorman · · focus · HN ↗
jp_gorman · · focus · HN ↗
epolanski · · focus · HN ↗
modeless · · focus · HN ↗
pmdr · · focus · HN ↗
samuelknight · · focus · HN ↗
mkaic · · focus · HN ↗
slekker · · focus · HN ↗
recursive · · focus · HN ↗
alirezaxdehghan · · focus · HN ↗
skerit · · focus · HN ↗
solarkraft · · focus · HN ↗
netruk44 · · focus · HN ↗
lynx97 · · focus · HN ↗
Aboutplants · · focus · HN ↗
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
<a href="https://www.engadget.com/2272106/openai-adds-dollar500-pro-subscription-nerfs-its-existing-dollar200-tier/" rel="nofollow">https://www.engadget.com/2272106/openai-adds-dollar500-pro-s...
Yikes
surgical_fire · · focus · HN ↗
The only way is for prices to go up. Way up.
mrtesthah · · focus · HN ↗
surgical_fire · · focus · HN ↗
machomaster · · focus · HN ↗
surgical_fire · · focus · HN ↗
The cost of providing the tokens for a heavy user (and let's be frank, the people paying $200 are likely heavy users) is many, many times more than the $200 recurring revenue they generate.
machomaster · · focus · HN ↗
Deepseek has low prices and despite that their profit margin at the beginning of this year was a whooping 82.9%. Since then, they have significantly raised prices.
You can actually check the approx. financials of OpenAI and Anthropic. The growth is insane.
There is no reason to believe why OAI/Anthropic wouldn't have a much better profit margin than DS, taking into account a much higher prices.
surgical_fire · · focus · HN ↗
> Deepseek has low prices and despite that their profit margin at the beginning of this year was a whooping 82.9%. Since then, they have significantly raised prices.
DeepSeek increased prices substantially not long ago. I find their profit margins hard to inspect considering I have very little idea what sort of environment they may get in China (from cheaper energy to government subsidies). I honestly doubt you have any insight here as well.
> You can actually check the approx. financials of OpenAI and Anthropic.
No you can't. They are not publicly traded, and they constantly and selectively leak bullshit metrics, from extremely unclear ARR, to extremely deceiving EBITDA. You willingly eat their bullshit and call me a picky eater in return.
> There is no reason to believe why OAI/Anthropic wouldn't have a much better profit margin than DS, taking into account a much higher prices.
I see no reason to believe (much less any actual evidence) that OAI or Anthropic have any path to profitability.
If inference (particularly for subscriptions) was in anyway as profitable as you claim today, they wouldn't need private investment rounds like crazy nor they would be desperate to offload this hot potato in an IPO.
82% margins lol. Are you telling me that if you created a machine that turns 1 dollar in 5 what you would do is dillute your ownership of the machine instead of using these fabulous profits to expand the business?
machomaster · · focus · HN ↗
It's clear that you are out of your depth when it comes to financials, business economics or a simple "what it takes to run a business".
You need money to make money. Growth strategy vs. self-financing strategy, pros and cons, when to do each. Critical chain. Limiting factor in infrastructure. Will not expand because this already goes over your head.
surgical_fire · · focus · HN ↗
I'm not the one out of my depth here.
Feel free to have the last word. I prefer to read idiocy in homeopatic doses.
moregrist · · focus · HN ↗
Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.
5555watch · · focus · HN ↗
Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.
With this in mind, it sounds like a fumble by OAI.
jpadkins · · focus · HN ↗
RussianCow · · focus · HN ↗
The vast majority of their revenue comes from large businesses buying for their teams, which are almost certainly not going to juggle lower tiers of different subscriptions to save a few bucks.
nananana9 · · focus · HN ↗
5555watch · · focus · HN ↗
Being grandfathered by OAI and happy is not the same as having both, and noticing "hmm maybe Claude is much better for my case, Ill suggest that to our manager"
RussianCow · · focus · HN ↗
TomGarden · · focus · HN ↗
Our VC-backed subscription days are numbered
Forgeties79 · · focus · HN ↗
m3kw9 · · focus · HN ↗
Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.
glaslong · · focus · HN ↗
onlyrealcuzzo · · focus · HN ↗
Well, the time it takes to compress frontier intelligence down to DeepSeek V4.1 Flash costs (basically too cheap to meter) is dropping, and the differential between the two is also dropping...
So... who cares?
honkycat · · focus · HN ↗
I can justify $200/mo but more than double is not appealing to me.
WinstonSmith84 · · focus · HN ↗
Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.
diffuse_l · · focus · HN ↗
spiderice · · focus · HN ↗
diffuse_l · · focus · HN ↗
cmrdporcupine · · focus · HN ↗
Yes, he was talking about safety, but IMHO they're likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.
I suspect we'll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.
enraged_camel · · focus · HN ↗
WinstonSmith84 · · focus · HN ↗
Come on .. this is barely released and you can already make that assessment?
And no, the $200 Anthropic plan is not significantly better than the $200 OpenAI plan, it's just the same Marketing non-sense and anybody shall now rather stick to the $100 plan of both of these provider if the monthly budget is $200. Anthropic doesn't have a Luna Max equivalent, and frankly Sol 6.1 is yet to be thoroughly tested.
[deleted] · · focus · HN ↗
[deleted]
MCArth · · focus · HN ↗
nostrebored · · focus · HN ↗
I think this is still true provided you're not using Astra.
machomaster · · focus · HN ↗
cromka · · focus · HN ↗
machomaster · · focus · HN ↗
An example. Let's assume that the work is evenly divided between days.
Imagine you want to work twice as much.
1. How efficiently can you use the reset credits if they would not reset the normal reset time?
Work with your normal weekly quota 3.5 days, press reset, work with new tokens for the rest of the week. Efficiency 100%.
2. With natural reset time going forward 7 days after each artificial reset.
You work for 3.5 days, press the reset, work for 3.5 days, wait for another 3.5 days for the natural reset, work for 3.5 days, press manual reset, work for 3.5 days, wait for 3.5 days... You can calculate the number for decreased efficiency yourself.
cromka · · focus · HN ↗
andriy_koval · · focus · HN ↗
the_duke · · focus · HN ↗
You have to do a lot of things in parallel.
InsideOutSanta · · focus · HN ↗
spiderice · · focus · HN ↗
Might want to hold off on canceling and continue to bleed them dry until the nerf hits
honkycat · · focus · HN ↗
torginus · · focus · HN ↗
adonese · · focus · HN ↗
scottLobster · · focus · HN ↗
cromka · · focus · HN ↗
glub · · focus · HN ↗
But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.
latentsea · · focus · HN ↗
glub · · focus · HN ↗
$500 for the old $200 is definitely a fumble.
latentsea · · focus · HN ↗
rrvsh · · focus · HN ↗
latentsea · · focus · HN ↗
RussianCow · · focus · HN ↗
Unless you're talking about buying enough hardware to run something like GLM 5.3, in which case the math just doesn't pencil out—the break even point is several years, and you're stuck with hardware that will be outdated well before then.
There are plenty of good reasons to use local models, but none of them are financial, at least for the vast majority of users.
latentsea · · focus · HN ↗
The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.
This way you're not actually at any disadvantage in terms of capability. You also don't need an advantage, you need to complete the tasks you care about. Eyes on the prize.
RTX 3090 came out a long time ago and it may be 'outdated' at this point but still banging like a champ for anyone who bought one and becoming increasingly more capable as new models unlock it's potential. Hardware hasn't changed much, but what it can do certainly has.
RussianCow · · focus · HN ↗
seizethecheese · · focus · HN ↗