GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Unofficial Hacker News client; not affiliated with Y Combinator.
the_duke · · focus · HN ↗
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
nxc18 · · focus · HN ↗
user43928 · · focus · HN ↗
The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.
I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.
nxc18 · · focus · HN ↗
I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.
holbrad · · focus · HN ↗
I haven't used it much yet, but I have much higher hopes for Sol 6.1, as it seems to be based off of a completely different base, it's not just a tune.
sebzim4500 · · focus · HN ↗
Astra is clearly far better than anything prior though, so I'm not sure what you mean really.
Hammershaft · · focus · HN ↗
Am I misinterpreting this, or did OpenAI clearly nerf GPT-6 Sol on the 23rd.
user43928 · · focus · HN ↗
The chart shows GPT-5.6 Sol and a surprisingly large drop in performance when the switched it over to GPT-6 Sol.
sigbottle · · focus · HN ↗
I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".
I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.
> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).
nxc18 · · focus · HN ↗
On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.
sigbottle · · focus · HN ↗
Yes, still running into this, but surprised about this
> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.
But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.
For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.
moshegramovsky · · focus · HN ↗
moshegramovsky · · focus · HN ↗
Here's a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That's a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn't. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn't follow existing API patterns.
jstummbillig · · focus · HN ↗
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
the_duke · · focus · HN ↗
phoghed · · focus · HN ↗
Codex itself seems to have a regression. You can see clearly the token use changing significantly coincides with a score drop
nicce · · focus · HN ↗
copperx · · focus · HN ↗
Marha01 · · focus · HN ↗
Eridrus · · focus · HN ↗
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
phoghed · · focus · HN ↗
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.
btbuildem · · focus · HN ↗
wkcheng · · focus · HN ↗
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
equinumerous · · focus · HN ↗
trentnix · · focus · HN ↗
chronogram · · focus · HN ↗
r0l1 · · focus · HN ↗
setnone · · focus · HN ↗
ozgung · · focus · HN ↗
bitexploder · · focus · HN ↗
NorthSouthNorth · · focus · HN ↗
sunaookami · · focus · HN ↗
stldev · · focus · HN ↗
For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.
For modeling and artwork, Astra has been great routinely outperforming Kimi.
This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.
I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.
rrvsh · · focus · HN ↗
I had to switch back to 5.6 Sol after trialling 6 Sol for like 3 days - I was getting insanely annoyed at how misaligned it is. Will try 6.1 but not very high hopes
keyle · · focus · HN ↗
I'll just leave this here: <a href="https://marginlab.ai/trackers/codex/" rel="nofollow">https://marginlab.ai/trackers/codex/
dannyw · · focus · HN ↗
And, is it really even an oligopoly anymore? Open weight models are incredibly competitive in every way; whether you want to use US providers, Chinese official providers, self host, etc.
moshegramovsky · · focus · HN ↗
I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30/40 minutes) to do simple things. As a result, Astra didn't get much done. It needs the same small implementation slices as GPT 5.5/others, but was much slower and didn't generate better results. (On a complex infra project/across a large codebase.)
It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn't seem to be better than 5.5 at most programming jobs.
I'm on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can't get coffee fast. Like I can't send an email fast.
OpenAI is making some excellent products for sure but I'm not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren't good at working autonomously on large codebase situations. Just because something compiles doesn't make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe/side app and then worked through the design there. Um, what? Not that it's invalid to do this but I actually have to test in the live codebase or I can't possibly say that something is working.
Just because you can, doesn't mean you should.
jrflo · · focus · HN ↗
skeptic_ai · · focus · HN ↗
soulofmischief · · focus · HN ↗
What was a pleasant and productive experience is becoming increasingly frustrating and draining.
beebmam · · focus · HN ↗
diego_sandoval · · focus · HN ↗
GPT 6 needs to be babysit, otherwise it starts doing ridiculous things.
4b11b4 · · focus · HN ↗
jsw97 · · focus · HN ↗
pampas · · focus · HN ↗
jeffybefffy519 · · focus · HN ↗
koyote · · focus · HN ↗
I've never seen such a large degradation in intelligence in a model until I tried out Sol 6 after having used 5.6 almost exclusively for several weeks.
twotwotwo · · focus · HN ↗
And, of course, GPT-6 came out as Anthropic fixed a bunch of stuff with their models -- faster (via fewer tokens, and TPS for Sonnet), easier to work with, better results, cheaper (via pricing and, again, fewer tokens). I don't know if the timing and the suddenness of the improvement on Anthropic's side sharpened the vibes comparison this round, but Internet opinion went pretty clearly to Anthropic.
FrontierCode's results make it look like Sol-6.1 may slot in well where you'd use Sonnet or Opus's low effort.
One thing I don't think any of this reflects is that many well-specified coding tasks, including the self-testing and doing research and tracing out dependencies and so on, aren't really bleeding-edge now: Luna-5.6 and small open models handle them fine. Stuff like "why is this box dropping connections?" or "here's a thing I want you to model/figure out" can benefit from bigger models. But far from everything does!
laurels-marts · · focus · HN ↗
I tried out fable 5.1 the day it was released and coming from gpt-5.6-sol I was truly mind blown (both in terms of code and prose it was generating - outputs I could finally enjoy reading and looking at).
Then when opus 5.5 came out, again same thing + far cheaper and faster.
I went from using OAI exclusively the entire year to a point now where i haven’t touched one of their models in at least a few weeks now.
I think OAI has lost the plot. OAI models simplify have no taste. And I don’t mean in front-end design way (although that too). They have no taste in how the model writes code, how it writes prose, how it writes in-line comments, how it writes documentation, or how it even picks variable names. There’s just no taste throughout.
Anthropic models are very thoughtful and have so much taste all around.
stasomatic · · focus · HN ↗
jp_gorman · · focus · HN ↗
jp_gorman · · focus · HN ↗