‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. phpnode · · focus · HN ↗
    What's driving the increase in release cadence here? We seem to get new models every week or so now, is this RSI?
    1. toasty228 · · focus · HN ↗
      Opus 5.5 is better than they anticipated, it's faster, smarter, cheaper. I'm about to change provider for claude and I'm not the only one
      1. copperx · · focus · HN ↗
        It feels like an updated 4.6. It's fantastic.
        1. RGS1811 · · focus · HN ↗
          I ran a battery of tests against a couple of simple prompts to check on thoroughness and verbosity of every available Opus, and 5.5 is a lot closer to 5 than people are letting on. 4.6 remains the best in terms of getting to the point and just doing what you ask. I had switched from 4.7 to 5.5 as my main claude model, but started running into the telltale over-interpretation issues of the 5 series, and have switched back. Something in their RL pipeline has made these models consistently worse IMO.
          1. copperx · · focus · HN ↗
            [delayed]
            1. RGS1811 · · focus · HN ↗
              My main grievance is that any gap in specificity in may statement of a task would lead Opus 5 to invent an interpretation to fill the gap, frequently creating lots of extra work for itself in the process, and often deviating from my intent. This would happen even for very simple things. I once asked Opus 5 to fix a failing unit test in CI (something pretty simple), and it went off on a 45 minute expedition (all in one turn), read a boatload of unnecessary files, massively overcomplicated the assignment, etc. It fixed the test but previous models would have handled this much more straightforwardly.

              A common form of this failure is the model picking up on random wordings from earlier in the session (e.g. some comment it made to me in the middle of a response, that I never explicitly endorsed) and then treating these as hard commitments. Or over-interpreting a specific word choice or clumsy phrasing as if it were a "load-bearing" constraint on the task.

              None of this clumsiness would be so problematic if the model didn't have such a strong drive toward autonomy. It's much like with people: there's no shame in not understanding what you're being asked to do, provided you ask clarifying questions. There's no shame in ignorance if it's wedded to curiosity. Benchmaxing has RLVRed curiosity and clarification straight out of these models. It sucks.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.