‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. gradus_ad · · focus · HN ↗
    Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
    1. mixdup · · focus · HN ↗
      Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability

      Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)

      1. sebzim4500 · · focus · HN ↗
        Is there anything that could happen that you wouldn't use as evidence that they are hitting a plateau?

        It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?

        1. djdjdkdkfk · · focus · HN ↗
          I am a lawyer not a coder, but for me the new models make the exact dumb mistakes they did in 2023. Everytime I come here I feel I'm in an alternate reality.
          1. RussianCow · · focus · HN ↗
            [delayed]
          2. robryan · · focus · HN ↗
            You get no better result from Opus 5.5 than Opus 4.1?
          3. epihelix · · focus · HN ↗
            IANAL, but it sounds as though law hallucinations are the final frontier. Sorry.

            But coding-wise, models keep getting better and cheaper. You can train for code correctness in a way you can't train for legal correctness, and you can test your code in an agentic loop in a way you can't test a legal opinion.

            Hence your alternative reality.

            (All that said, 2023 was GPT-4 territory. GPT-4o wasn't released until 2024. No matter what question you're asking, I struggle to believe you wouldn't notice the difference between GPT-4 and the current frontier model set. You can download and run any number of sub-27B local models that will be better than GPT-4. The pace of change in this field really has been insane.)

          4. rspeele · · focus · HN ↗
            Both the rate of improvement and the actual effectiveness of LLMs is much higher in coding than in other fields. An LLM+coding harness can try something wrong and get rapidly, harmlessly caught by a compiler or by a unit test. Then it usually gets it right on the 2nd or 3rd attempt. This all happens before the user who requested the change is told it's done. And a lot of it goes right back into the training dataset for the next model release.

            In law if a model makes a mistake it pretty much takes a human to catch it, which is a vastly slower, riskier, and more frustrating feedback loop. So it doesn't surprise me they can stay dumb in that field while getting massively more capable in coding with multiple "step change" releases in the past 365 days.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.