‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. the_duke · · focus · HN ↗
    The GPT 6 release was ... not great.

    Sol 6 was so bad that I switched over to Opus 5.5 exclusively.

    Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.

    Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.

    I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.

    (Note: this is after preferring and shilling Codex/OpenAI models for the last half year)

    1. nxc18 · · focus · HN ↗
      How does this jive with the exponential growth claims? Theoretically sol models are better than the 4 series models I was using at the beginning of the year, but in practice the results don’t seem to be much better. They always nerf the models over the course of the release so it _looks_ like the next version is better but I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
      1. user43928 · · focus · HN ↗
        They never nerfed any model after release.

        The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.

        I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.

        1. nxc18 · · focus · HN ↗
          I am comparing to GPT-5.3 and 5.2, and I perceive that things have not been noticeably better since then. I also know that I can predict new model releases with high accuracy when my coding agent suddenly becomes regard-level at following instructions and completing simple tasks. This is how I knew 6.0 was about to be released - 5.6 suddenly got unusably bad.

          I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.

          1. holbrad · · focus · HN ↗
            I've heard very little positive press around Sol 6, with a ton of people preferring Sol 5.6 instead.

            I haven't used it much yet, but I have much higher hopes for Sol 6.1, as it seems to be based off of a completely different base, it's not just a tune.

          2. sebzim4500 · · focus · HN ↗
            I never used sol-6.0 in part because everyone kept talking about how bad it was.

            Astra is clearly far better than anything prior though, so I'm not sure what you mean really.

        2. Hammershaft · · focus · HN ↗
          <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;

          Am I misinterpreting this, or did OpenAI clearly nerf GPT-6 Sol on the 23rd.

          1. user43928 · · focus · HN ↗
            It released on that date, the data before is a different model.

            The chart shows GPT-5.6 Sol and a surprisingly large drop in performance when the switched it over to GPT-6 Sol.

      2. sigbottle · · focus · HN ↗
        How large of codebases are you working on? The models have gotten good enough to 1 shot stupid &quot;trivial&quot; throwaway integration projects with 0 handholding (was having RL&#x27;d garbage in late 2025), and I&#x27;m actually enjoying designing bounded greenfield personal software from scratch with Astra, in my experience. It&#x27;s quite slow - 2 weeks of credits and constant talking and back and forth with Astra, but it doesn&#x27;t feel annoying to talk to and is like an intelligent colleague maybe 70% of the time? Which is great. Just push back when it&#x27;s dumb.

        I&#x27;m by no means an AI booster, but given 2022 - 2026 progress I&#x27;d say it&#x27;s &quot;exponential&quot; in the sense of, &quot;holy shit, every year I can do more and more genuinely different things&quot;, not &quot;RSI mind reading intelligence can do anything is here&quot;.

        I don&#x27;t think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.

        &gt; I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling&#x2F;harness improvements.

        Even if that were the case, I&#x27;d say that it&#x27;s improved in practice. And just from a philosophy perspective, if you&#x27;re trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).

        1. nxc18 · · focus · HN ↗
          It’s 50&#x2F;50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).

          On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.

          5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.

          1. sigbottle · · focus · HN ↗
            &gt; It’s 50&#x2F;50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).

            Yes, still running into this, but surprised about this

            &gt; On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.

            I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just &quot;grasp&quot; the right level of &quot;here is the essence of what we need&quot; versus &quot;these are all the small impl details&quot;. But idk I feel like Astra&#x27;s the first model in quite a while that I don&#x27;t feel genuinely annoyed at handholding a toddler with a PhD.

            But I totally believe you on the 50&#x2F;50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of &quot;make user retry in this case&quot;, it silently built an extremely elaborate recovery state machine w&#x2F;o looking. These pathologies by no means gone, and I&#x27;m still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra&#x27;s gonna do this kind of RL slop failure mode.

            For my use cases personally though, it&#x27;s been better and better. I can&#x27;t use AI at work, so you have much harier edge cases than I do, but still.

          2. moshegramovsky · · focus · HN ↗
            This is 100% absolutely my experience as well. Especially the needless abstractions and endless rounds of corrections. That was literally my entire last week of work.
        2. moshegramovsky · · focus · HN ↗
          I work on a very large code base (millions of LOC) and I&#x27;ve had lackluster results with autonomous work and 1 shotting. AI is definitely fantastic at working on many programming problems but I am not seeing amazing results at refactoring. In fact, I am seeing very poor results, even with Astra, even with extensive planning docs. All the recent models I&#x27;ve used can definitely get that refactor done, but not autonomously. It needs to be small slices. I&#x27;ve yet to see it 1 shot anything really complicated.

          Here&#x27;s a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That&#x27;s a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn&#x27;t. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn&#x27;t follow existing API patterns.

    2. jstummbillig · · focus · HN ↗
      Eh. What? Is this common sentiment?

      I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?

      (Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)

      1. the_duke · · focus · HN ↗
        On r&#x2F;codex the sentiment seems to be quite wide-spread.
        1. phoghed · · focus · HN ↗
          <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;

          Codex itself seems to have a regression. You can see clearly the token use changing significantly coincides with a score drop

      2. nicce · · focus · HN ↗
        When GPT 6 Sol &amp; Luna were released, everything went down. I have been running Sol at max thinking and it is about the same as old Luna with max thinking, give or take. Sometimes feeling even dumber. I can&#x27;t trust it to do anything big alone anymore without babysitting.
      3. copperx · · focus · HN ↗
        Opus 5.5 is so good that I don&#x27;t want it to be replaced anytime soon. Stop training models, Anthropic, and just serve this thing without regressions for a year or three, can you?
        1. Marha01 · · focus · HN ↗
          They should etch it into an ASIC. The first model worthy of that honor.
      4. Eridrus · · focus · HN ↗
        Sol 6 definitely feels kind of dumb and worse than 5.6

        Astra seems better though.

        Showing one potentially saturated benchmark doesn&#x27;t necessarily fill me with a lot of confidence in the coding results.

      5. phoghed · · focus · HN ↗
        In my experience, no. There’s no way to know though. The whole conversation and industry are a combo of benchmaxing, faith, and mysticism.

        Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.

    3. btbuildem · · focus · HN ↗
      That mirrors how disappointing Opus 5 and Fable were, for anything beyond one-shotted tasks or shiny demos. Maybe OAI is just a step behind Anthropic? Opus 5.5 seems like the real deal again, consistent good results on large, complex codebases.
    4. wkcheng · · focus · HN ↗
      I agree, and I haven&#x27;t seen other people mention this! The benchmarks for GPT 6 Sol are great, but realistically it does not seem better than 5.6 Sol. 6-Sol is noticeably worse for code reviews (worse than Deepseek 4.1 flash), has implementation issues (requires more rounds of code reviews and fixes to get to a serviceable state). Opus 5.5 is much much better.

      I&#x27;ve implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.

      If 6.1 Sol has actually matched Opus 5.5, I&#x27;d be very happy. However, benchmarks and real usage don&#x27;t seem to agree in my own tests. So we&#x27;ll have to see.

      1. equinumerous · · focus · HN ↗
        If the benchmarks show better performance, but a consensus of experienced software engineers establishes that the model is worse on coding performance... well, the benchmarks don&#x27;t mean much, do they? It seems like we need much more comprehensive and better benchmarks. And of course, I don&#x27;t think benchmarks yet capture the &quot;human&quot; factor - does a human think a bit of code is logical and maintainable? I often find that these models produce a bit of code, but it is much more convoluted than it needs to be. It makes perfect sense given that these things are code generators, that they generate a lot of code. But quantity of code does not mean code quality, and code quality tends to matter when you read code much more than you write it.
    5. trentnix · · focus · HN ↗
      That&#x27;s not been my experience. My experience with Astra (I use it at home writing Go and C) for coding has been fantastic. Opus 5.5 (I use it for work writing C#) seems faster than Opus 5, but it doesn&#x27;t seem demonstrably better to my eyes and is still prone to word vomit.
      1. chronogram · · focus · HN ↗
        Same here. Astra has been the best thing I&#x27;ve seen. Astra on Low has been my favourite thing so far. Higher levels just mean more cruft, not useful.
      2. r0l1 · · focus · HN ↗
        Made the opposite experience. Astra was not good in writing go and c++ code. Had multiple OpenAi and Claude subscriptions and all our coworkers agreed. Switched back to Claude and the experience is so much better. Not vibe coding, but assisted coding with immediate feedback.
    6. setnone · · focus · HN ↗
      yeah i can relate, sol 6 is definitely dumber than 5.6, lazier too, i hope it&#x27;s just roll out pains
    7. ozgung · · focus · HN ↗
      Maybe OpenAI was the only one pacing the frontier.
    8. bitexploder · · focus · HN ↗
      I have likewise not been impressed with Astra 6 for most things. It is good, but Opus 5.5 seems just as good or better and I have had Opus 5.5 workers just... hammering since release and cannot spend all of my quota yet.
    9. NorthSouthNorth · · focus · HN ↗
      I shilled so hard to a friend that he actually swapped decided to swap over to Codex. I feel a bit guilty now lol (tbh Astra is a great model, but 5.5 is just brilliant).
    10. sunaookami · · focus · HN ↗
      gpt-6-luna is terrible. It leaks tool calls and markers in the output like crazy, there is definitely something wrong here. gpt-5.6-terra works fine. Also, gpt-6-luna was sneakily added to the 1 mio free tokens group instead of 10 mio. like gpt-5.6-luna: <a href="https:&#x2F;&#x2F;help.openai.com&#x2F;en&#x2F;articles&#x2F;10306912-sharing-feedback-evaluation-and-fine-tuning-data-and-api-inputs-and-outputs-with-openai#:~:text=What%20models%20are%20included%20in%20this%20offer" rel="nofollow">https:&#x2F;&#x2F;help.openai.com&#x2F;en&#x2F;articles&#x2F;10306912-sharing-feedbac...
    11. stldev · · focus · HN ↗
      My experience as well.

      For coding specifically, I&#x27;ve found 5.6-Sol &gt; 6.0 Sol &gt; Astra.

      For modeling and artwork, Astra has been great routinely outperforming Kimi.

      This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.

      I can&#x27;t wait for technology to catch up to a point where we can rid ourselves of this oligopoly.

      1. rrvsh · · focus · HN ↗
        Hard agree

        I had to switch back to 5.6 Sol after trialling 6 Sol for like 3 days - I was getting insanely annoyed at how misaligned it is. Will try 6.1 but not very high hopes

      2. keyle · · focus · HN ↗
        I&#x27;d even go one more, 5.4 was great until 5.5, which was a rug pull.

        I&#x27;ll just leave this here: <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;

      3. dannyw · · focus · HN ↗
        Adaptive thinking was a good idea though. The old method of manually specifying how many thinking tokens you wanted as budget was just silly. The rollout might not have been great, but the change is good.

        And, is it really even an oligopoly anymore? Open weight models are incredibly competitive in every way; whether you want to use US providers, Chinese official providers, self host, etc.

    12. moshegramovsky · · focus · HN ↗
      100% hard agree.

      I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30&#x2F;40 minutes) to do simple things. As a result, Astra didn&#x27;t get much done. It needs the same small implementation slices as GPT 5.5&#x2F;others, but was much slower and didn&#x27;t generate better results. (On a complex infra project&#x2F;across a large codebase.)

      It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn&#x27;t seem to be better than 5.5 at most programming jobs.

      I&#x27;m on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can&#x27;t get coffee fast. Like I can&#x27;t send an email fast.

      OpenAI is making some excellent products for sure but I&#x27;m not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren&#x27;t good at working autonomously on large codebase situations. Just because something compiles doesn&#x27;t make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe&#x2F;side app and then worked through the design there. Um, what? Not that it&#x27;s invalid to do this but I actually have to test in the live codebase or I can&#x27;t possibly say that something is working.

      Just because you can, doesn&#x27;t mean you should.

    13. jrflo · · focus · HN ↗
      I&#x27;m in the same boat, I&#x27;ll give 6.1 a shot but I&#x27;ll probably hop over to Anthropic now that the $200 tier has equivalent weekly usage between the two of them.
      1. skeptic_ai · · focus · HN ↗
        Anthropic 20x plan only refers to 5h interval. Not the weekly quota. Very sketchy
    14. soulofmischief · · focus · HN ↗
      I have had the same exact experience. I feel like I&#x27;m working with 5.3 again. It is alarming how degraded the experience has become over the last month.

      What was a pleasant and productive experience is becoming increasingly frustrating and draining.

    15. beebmam · · focus · HN ↗
      gpt-5.6-sol is significantly better than gpt-6-sol. Not impressed with this new line.
      1. diego_sandoval · · focus · HN ↗
        Agree.

        GPT 6 needs to be babysit, otherwise it starts doing ridiculous things.

      2. 4b11b4 · · focus · HN ↗
        Didn&#x27;t even bother trying it yet
    16. jsw97 · · focus · HN ↗
      After seeing a number of hit or miss releases from both OpenAI and Anthropic my default is to stay put on what I’m using and then free ride on discerning eager adopters by reading their reviews. (Thanks!) Still on sol 5.6 with an occasional advice from Astra. Also I feel like I kind of get used to the models but maybe that’s just my imagination.
    17. pampas · · focus · HN ↗
      That&#x27;s my experience too. GPT-6 Sol tends to rabbit hole and over engineer things.
    18. jeffybefffy519 · · focus · HN ↗
      Its almost like the &quot;frontier&quot; is a load of marketing bullshit and we should ignore it....
    19. koyote · · focus · HN ↗
      I think the fact that Sol 6 appeared higher than Sonnet 4 on benchmarks shows that benchmarks are completely rubbish and useless.

      I&#x27;ve never seen such a large degradation in intelligence in a model until I tried out Sol 6 after having used 5.6 almost exclusively for several weeks.

    20. twotwotwo · · focus · HN ↗
      I am always uncertain about impressions, but mine agree with this. I liked Luna 5.6 on Amazon Bedrock (which got &gt;100 tps) for doing well-specced tasks fast. 6 seems to both be served slower by Bedrock and may spend more turns&#x2F;tokens to get to the same place, so...not as fun.

      And, of course, GPT-6 came out as Anthropic fixed a bunch of stuff with their models -- faster (via fewer tokens, and TPS for Sonnet), easier to work with, better results, cheaper (via pricing and, again, fewer tokens). I don&#x27;t know if the timing and the suddenness of the improvement on Anthropic&#x27;s side sharpened the vibes comparison this round, but Internet opinion went pretty clearly to Anthropic.

      FrontierCode&#x27;s results make it look like Sol-6.1 may slot in well where you&#x27;d use Sonnet or Opus&#x27;s low effort.

      One thing I don&#x27;t think any of this reflects is that many well-specified coding tasks, including the self-testing and doing research and tracing out dependencies and so on, aren&#x27;t really bleeding-edge now: Luna-5.6 and small open models handle them fine. Stuff like &quot;why is this box dropping connections?&quot; or &quot;here&#x27;s a thing I want you to model&#x2F;figure out&quot; can benefit from bigger models. But far from everything does!

    21. laurels-marts · · focus · HN ↗
      100% in agreement. I pay for OAI sub and also use Codex exclusively at work for the past 8 months.

      I tried out fable 5.1 the day it was released and coming from gpt-5.6-sol I was truly mind blown (both in terms of code and prose it was generating - outputs I could finally enjoy reading and looking at).

      Then when opus 5.5 came out, again same thing + far cheaper and faster.

      I went from using OAI exclusively the entire year to a point now where i haven’t touched one of their models in at least a few weeks now.

      I think OAI has lost the plot. OAI models simplify have no taste. And I don’t mean in front-end design way (although that too). They have no taste in how the model writes code, how it writes prose, how it writes in-line comments, how it writes documentation, or how it even picks variable names. There’s just no taste throughout.

      Anthropic models are very thoughtful and have so much taste all around.

      1. stasomatic · · focus · HN ↗
        I cancelled Claude because of its thoughtful prose. I prefer one liner responses from OAI models.
    22. jp_gorman · · focus · HN ↗
      yup - It was darn rude (as the Aussies would say)... Was taking forever to do anything and it made a mess of a lot of resourcing I was working on for som Dev Ops work I was doing on an AI project. Switched back to 5.6 sol and it immediately noted all the mess of worktrees and PRs it had left all over the place.
    23. jp_gorman · · focus · HN ↗
      Fully agree - had to drop to 5.6 sol as 6.1 sol was atrain wreck in my project leaving dead boddies everywhere... 5.6 sol spotted it all the minute it looked. 6.1 sol was noticably major slow down also.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.