‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. gradus_ad · · focus · HN ↗
    Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
    1. mixdup · · focus · HN ↗
      Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability

      Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)

      1. semiquaver · · focus · HN ↗
        What universe do you live in that you can look at the past six months and see anything like a plateau in capability?

        Edit: removed a comment that was uncharitable and rude, for which I apologize.

        1. ActionHank · · focus · HN ↗
          Have we honestly seen that great a leap in the last 6 months, or just better application of what we had 6 months before that.

          We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.

          1. CuriouslyC · · focus · HN ↗
            The difference between 6 months ago frontier and now frontier in 3d modelling, graphics and video editing is night and day.
            1. ActionHank · · focus · HN ↗
              Just because there are new capabilities, doesn't mean they've pushed passed the plateau, they've just expanded where the previous solutions work.

              We've gone from 80% in some places to 80% in some more places.

              1. CuriouslyC · · focus · HN ↗
                > We've gone from 80% in some places to 80% in some more places.

                Any area that is verifiable will trend inexorably towards 100% over time. In unverifiable areas, it'll always be "80%" because the ubiquity of "AI" style erodes its value, and ">80%" for unverifiable things involves fashion, cachet and "vibes" that humans will probably never knowingly let it have.

                1. ActionHank · · focus · HN ↗
                  Just checked your website, you really drank all the koolaid huh?
              2. coderenegade · · focus · HN ↗
                I don't think the math community thinks there's been a plateau. The models have gone from being a joke to being able to crank out proofs to research grade problems.
          2. usef- · · focus · HN ↗
            Statements like these remind me how much each of us are in our own unique information bubbles. People seemed very excited about analysing the differences of each model on my feeds.

            Have you used recent ones on any large projects or coding issues? They've improved tremendously lately.

        2. mixdup · · focus · HN ↗
          Not that they've hit it but that they are approaching it. The time to panic and steer the narrative is before you hit the iceberg, not after
          1. helloplanets · · focus · HN ↗
            The whole "pace the frontier" thing originated from Anthropic.

            With Opus 5.5, it doesn't seem like model improvement is plateuaing. And Fable 5.5 will likely be dropped this or next week.

            You can only imagine what they've got going internally.

        3. phoghed · · focus · HN ↗
          People have been saying this since GPT-4.
        4. arctic-true · · focus · HN ↗
          Most of the impressive accomplishments we’ve seen in the last few months have been the result of huge agent swarms working together and brute-forcing solutions, not massive leaps in intelligence from standalone models. That is still an improvement in the usefulness and power of the technology, but it is NOT evidence that model intelligence is increasing faster than before.
          1. famouswaffles · · focus · HN ↗
            I don't have any access to any agent swarms (and neither do most) and i still think the models have obviously improved massively in standalone intelligence. Of course they have, agent swarms are not magic. You can swarm all you want around GPT-4 era models and you'll get nowhere. And i've never seen the term 'brute-force' more abused than these LLM discussions. Basically none of the results have been brute force.
            1. semiquaver · · focus · HN ↗
              Agreed. You can’t “brute force” reality, which has an infinitely large state space. A million monkeys won’t write Shakespeare and all that
          2. CamperBob2 · · focus · HN ↗
            "This machine-intelligence stuff is overrated, they are just using <insert particular machine-intelligence technique here>" isn't the resounding verdict it may have sounded like when you typed it.
          3. usef- · · focus · HN ↗
            Do you use them? I feel like you're talking about the news headlines. But in daily, direct, individual usage (one model at a time) I've found they've improved tremendously on the last few months.
            1. arctic-true · · focus · HN ↗
              Yes, I do. They’ve definitely gotten better, I wouldn’t dispute that. But we haven’t gone from useless toys to Skynet in the last 6 months, progress is slower and steadier than that.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.