‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. the_duke · · focus · HN ↗
    The GPT 6 release was ... not great.

    Sol 6 was so bad that I switched over to Opus 5.5 exclusively.

    Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.

    Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.

    I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.

    (Note: this is after preferring and shilling Codex/OpenAI models for the last half year)

    1. wkcheng · · focus · HN ↗
      I agree, and I haven't seen other people mention this! The benchmarks for GPT 6 Sol are great, but realistically it does not seem better than 5.6 Sol. 6-Sol is noticeably worse for code reviews (worse than Deepseek 4.1 flash), has implementation issues (requires more rounds of code reviews and fixes to get to a serviceable state). Opus 5.5 is much much better.

      I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.

      If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.

      1. equinumerous · · focus · HN ↗
        If the benchmarks show better performance, but a consensus of experienced software engineers establishes that the model is worse on coding performance... well, the benchmarks don't mean much, do they? It seems like we need much more comprehensive and better benchmarks. And of course, I don't think benchmarks yet capture the "human" factor - does a human think a bit of code is logical and maintainable? I often find that these models produce a bit of code, but it is much more convoluted than it needs to be. It makes perfect sense given that these things are code generators, that they generate a lot of code. But quantity of code does not mean code quality, and code quality tends to matter when you read code much more than you write it.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.