‹ BackHN Continuity

Thread

GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence

80 points · 102 comments · theanonymousone

  1. madsgarff · · focus · HN ↗
    It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode. I use claude, and I wanna build a feeling for what high, medium, etc. actually gives me. So far their comparisons, and having tried several different models for my work, has given me a feel of what 50 intelligence actually is. And I believe it would be be of even greater value to get a feel inside the single model I actually use, as most people do, because not many, I believe, switch heavily between models when working. I understand that the cost here is greater but the model provivders should obviously give you free access, because of the great work you are doing.
    1. rrr_oh_man · · focus · HN ↗
      > I wanna build a feeling for what high, medium, etc. actually gives me

      Nothing, really. It's like oversampling your data set. You usually get a much better overall baseline performance if you use the default setting.

      1. lumzell · · focus · HN ↗
        In my case it actually made a difference. I tested it on my own code review set with Gemini Flash. On low it found fewer bugs than on default around 90% versus 97% but was about three times faster and a lot cheaper. On high it actually found everything in the hardest case but took almost three minutes per call. so I think the difference is real you only see it if you run the same fixed cases a few times and not by feel. And yes I used Gemini rather than the GPT Sol that is mentioned here would be interesting to see the same tests on Sol.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.