‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. phpnode · · focus · HN ↗
    What's driving the increase in release cadence here? We seem to get new models every week or so now, is this RSI?
    1. jesse_dot_id · · focus · HN ↗
      No.
      1. anotha_one · · focus · HN ↗

        [dead]

    2. lxgr · · focus · HN ↗
      Wanting to have the newer model than the competitor, presumably.
      1. dandellion · · focus · HN ↗
        The old "the bigger number is better", GPT announces model 6.1, the obvious thing to do next is to announce Gemini 27, and after that Claudé 3000, then a flute album.
        1. lxgr · · focus · HN ↗
          We swear, We Really Wanted To Make An "ASI" Model But This Is Literally The Way The Weights Dragged Us This Time
    3. sharpshadow · · focus · HN ↗
      Response to DeepSeek’s technical paper and competition.
      1. wg0 · · focus · HN ↗
        What's that in summary?
        1. Wheen · · focus · HN ↗
          Not the person you're replying to, but judging by the emphasis on the cost of cached input tokens in the OP article, I'd guess it has to do with DeepSeek v4.1's KV cache efficiency. It uses <1000 bytes per token, so they're able to get 1M token context in under a GB.

          Edit: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;deepseek-ai&#x2F;DeepSeek-V4.1-Flash&#x2F;blob&#x2F;main&#x2F;DeepSeek_V41_Tech_Report.pdf" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;deepseek-ai&#x2F;DeepSeek-V4.1-Flash&#x2F;blob&#x2F;...

          1. ChromeUltron · · focus · HN ↗
            just goes to show that OpenAI in fact did not innovate on a single thing for the better part of a year (one could argue two) and instead keeps immitating what it sees doing others successfully with the tech, all in a very transparent attempt to get people lubed up for their IPO.
      2. LPisGood · · focus · HN ↗
        Which paper are you referring to?
    4. mckirk · · focus · HN ↗
      No, we&#x27;re pacing ourselves to have the time to evaluate the impact each new model could have, obviously.
    5. system2 · · focus · HN ↗
      Chinese model pressure. Many of my SWE friends switched to Chinese models. I also use QWEN and GLM for many of the api requiring projects and dropped OpenAI and Anthropic. The only reason was the cost.

      EDIT: I love getting downvoted by openai and anthropic employees or their bots.

      1. wg0 · · focus · HN ↗
        I can&#x27;t recommend Chinese models enough. My personal favorite is DeepSeek v4.1 Flash but I have tried Qwen 3.8, Kimi 3 and GLM 5.3 which are equally impressive but DeepSeek is the cheapest and fastest regularly hitting 270 token per second.

        And yeah I have worked with Anthropic and OpenAI models, they&#x27;re good but they cost a fortune while Chinese models are already really good at a fraction of the cost.

        1. copperx · · focus · HN ↗
          [delayed]
        2. andybak · · focus · HN ↗
          DeepSeek v4.1 Flash is fascinating and uneven. It&#x27;s way too chatty in OpenCode to be a collaboration partner. I tried dsh-tui which feels comparable to the codex&#x2F;claude tui&#x27;s and it&#x27;s usable. but it seems to be &quot;brilliant and yet stupid&quot; in a way I can&#x27;t quite put my finger on. I&#x27;ve got too much real work to get done to dig into it so until the big boys price me out of the market I&#x27;m back to my $100&#x2F;month deal.
      2. thraway3837 · · focus · HN ↗
        I keep hearing about these Chinese models, but what exactly are you doing with the models and coding? I have a need to fully write code with full tool calling capabilities. Not just methods or functions. I want to be able to prompt a feature and it makes the JIRA ticket, and fully implements it and makes a PR. I don&#x27;t want to babysit it or even read the code. Once it creates the PR, I want it to monitor it for any comments fro Copilot&#x2F;security review and then fix it as necessary.

        Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?

        1. senderista · · focus · HN ↗
          Sounds like your dream workflow could replace you with a PM?
        2. white_dragon88 · · focus · HN ↗

          [dead]

        3. system2 · · focus · HN ↗
          I am using the API only with them for now. But what you describe is nothing compared to Qwen or Mimo. These models are more capable than Opus in general and at a fraction of Opus&#x27;s cost.
    6. jonatron · · focus · HN ↗
      Probably just the singularity, no big deal
      1. anotha_one · · focus · HN ↗

        [dead]

    7. SwabbyNat74 · · focus · HN ↗
      Its a news cycle more than anything, and its ONLY going to get much, much worse. Daily releases, or multiple daily, 30-45, by EOY. Welcome to RSI!
    8. tjwebbnorfolk · · focus · HN ↗
      Competition
    9. colpabar · · focus · HN ↗
      What I don&#x27;t understand is how much people have to say about every single one. Aren&#x27;t we at the diminishing returns stage yet? Is there really that much to discuss?
      1. infamouscow · · focus · HN ↗
        If you look closely at various benchmarks, you&#x27;ll see that often models will improve in certain areas while regressing in others. It suggests we&#x27;re already at the point of diminishing returns.
    10. Aboutplants · · focus · HN ↗
      I do wonder if people switch back and forth between primary models (GPTvsClaude) that it may be a better idea to simply keep releasing updates as soon as possible in order to keep users from bouncing back and forth.
      1. scrollop · · focus · HN ↗
        Probably one of the factors. Signed up to openai pro a few days ago, deciding between openai and anthropic, then sonnet 5.5 was released and am wondering whether I made a mistake.

        Luckily it&#x27;s not a mistake as now we have access to . . . dots.

        (and sol 6.1, it seems)

      2. sockaddr · · focus · HN ↗
        This is it.

        It&#x27;s because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor&#x27;s version bump is the logical thing to do. It has nothing to do with RSI.

        1. vividfrier · · focus · HN ↗

          [dead]

      3. geeky4qwerty · · focus · HN ↗
        jokes on me, I pay for all the subscriptions.
      4. pythonaut_16 · · focus · HN ↗
        Maybe process maturity too.

        Like think about a software org with good CI&#x2F;CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.

        As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.

      5. killingtime74 · · focus · HN ↗
        Of course they do. The real money makers are not subscription users, but the API users and you can just switch with the model selector.
    11. mattnewton · · focus · HN ↗
      Anthropic’s IPO?
    12. franzcoughka · · focus · HN ↗

      [dead]

    13. motoboi · · focus · HN ↗
      New models are distill from the actual unrelease frontier models. They are just giving us better checkpoints.
    14. mynameisjonny_ · · focus · HN ↗
      The initial response to 6 Sol was bad, and Opus 5.5 was definitely winning the public vibes war. Makes sense to rush something out
    15. agluszak · · focus · HN ↗
      They&#x27;re releasing Sol 6.1 because 1. Astra 6.1 got postponed 2. Sol 6 is shitty 3. They have to release _something_ in response to Opus 5.5
    16. orbital-decay · · focus · HN ↗
      Versions is marketing, snapshots&#x2F;minor variations are easy and the number must go up. Release timing is another OAI&#x27;s marketing tactic.

      &gt;RSI

      Recursive improvement doesn&#x27;t imply increased rate, another word for it is &quot;iterative&quot; but this probably sounds too boring to some people.

    17. jchw · · focus · HN ↗
      It is the only way to reduce prices while making it look like a good thing.
    18. esafak · · focus · HN ↗
      Productivity is increasing as models get smarter; we are ascending the singularity. I&#x27;m serious.
    19. toasty228 · · focus · HN ↗
      Opus 5.5 is better than they anticipated, it&#x27;s faster, smarter, cheaper. I&#x27;m about to change provider for claude and I&#x27;m not the only one
      1. copperx · · focus · HN ↗
        It feels like an updated 4.6. It&#x27;s fantastic.
        1. RGS1811 · · focus · HN ↗
          I ran a battery of tests against a couple of simple prompts to check on thoroughness and verbosity of every available Opus, and 5.5 is a lot closer to 5 than people are letting on. 4.6 remains the best in terms of getting to the point and just doing what you ask. I had switched from 4.7 to 5.5 as my main claude model, but started running into the telltale over-interpretation issues of the 5 series, and have switched back. Something in their RL pipeline has made these models consistently worse IMO.
          1. copperx · · focus · HN ↗
            [delayed]
            1. RGS1811 · · focus · HN ↗
              My main grievance is that any gap in specificity in may statement of a task would lead Opus 5 to invent an interpretation to fill the gap, frequently creating lots of extra work for itself in the process, and often deviating from my intent. This would happen even for very simple things. I once asked Opus 5 to fix a failing unit test in CI (something pretty simple), and it went off on a 45 minute expedition (all in one turn), read a boatload of unnecessary files, massively overcomplicated the assignment, etc. It fixed the test but previous models would have handled this much more straightforwardly.

              A common form of this failure is the model picking up on random wordings from earlier in the session (e.g. some comment it made to me in the middle of a response, that I never explicitly endorsed) and then treating these as hard commitments. Or over-interpreting a specific word choice or clumsy phrasing as if it were a &quot;load-bearing&quot; constraint on the task.

              None of this clumsiness would be so problematic if the model didn&#x27;t have such a strong drive toward autonomy. It&#x27;s much like with people: there&#x27;s no shame in not understanding what you&#x27;re being asked to do, provided you ask clarifying questions. There&#x27;s no shame in ignorance if it&#x27;s wedded to curiosity. Benchmaxing has RLVRed curiosity and clarification straight out of these models. It sucks.

      2. copperx · · focus · HN ↗
        &gt; I&#x27;m not the only one

        See, that&#x27;s an&#x2F;the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.

    20. denysvitali · · focus · HN ↗
      They&#x27;re pacing the frontier
      1. blmarket · · focus · HN ↗
        and seems like they&#x27;re claiming Sol&#x2F;Opus are not frontier (and only Astra&#x2F;Fable are)
    21. az226 · · focus · HN ↗
      Mature training pipelines, plus ever expanding RL datasets of increased quality, and mega GPU clusters to finish training in a few weeks. Automated safety and reliability testing.
    22. MisterMunchkin · · focus · HN ↗
      Both labs are spying on each other and they get jelly when the other is releasing a new model, so they have to ship something at the same time so they don’t look bad.
      1. dannyw · · focus · HN ↗
        Both labs are spying? Employees hang out at the same bars, have overlapping social circles, etc. Alcohol does what alcohol does.
    23. ChromeUltron · · focus · HN ↗
      no patrick, m̶a̶y̶o̶n̶n̶a̶i̶s̶e̶ a point release of the slopbot is NOT a̶n̶ i̶n̶s̶t̶r̶u̶m̶e̶n̶t̶ RSI
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.