‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. the_duke · · focus · HN ↗
    The GPT 6 release was ... not great.

    Sol 6 was so bad that I switched over to Opus 5.5 exclusively.

    Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.

    Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.

    I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.

    (Note: this is after preferring and shilling Codex/OpenAI models for the last half year)

    1. twotwotwo · · focus · HN ↗
      I am always uncertain about impressions, but mine agree with this. I liked Luna 5.6 on Amazon Bedrock (which got >100 tps) for doing well-specced tasks fast. 6 seems to both be served slower by Bedrock and may spend more turns/tokens to get to the same place, so...not as fun.

      And, of course, GPT-6 came out as Anthropic fixed a bunch of stuff with their models -- faster (via fewer tokens, and TPS for Sonnet), easier to work with, better results, cheaper (via pricing and, again, fewer tokens). I don't know if the timing and the suddenness of the improvement on Anthropic's side sharpened the vibes comparison this round, but Internet opinion went pretty clearly to Anthropic.

      FrontierCode's results make it look like Sol-6.1 may slot in well where you'd use Sonnet or Opus's low effort.

      One thing I don't think any of this reflects is that many well-specified coding tasks, including the self-testing and doing research and tracing out dependencies and so on, aren't really bleeding-edge now: Luna-5.6 and small open models handle them fine. Stuff like "why is this box dropping connections?" or "here's a thing I want you to model/figure out" can benefit from bigger models. But far from everything does!

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.