‹ BackHN Continuity

Thread

GPT-6 Sol and Luna

1779 points · 855 comments · OfficialTurkey

  1. m_fayer · · focus · HN ↗
    I've been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It's the first model I've gotten attached to. I'm concerned that whatever model supercedes it, while technically better, just won't feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.
    1. redox99 · · focus · HN ↗
      Same. In fact I found 6 Astra to be a downgrade in situations where I didn't need the extra intelligence.
      1. cmrdporcupine · · focus · HN ↗
        Yeah.

        Astra was/is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague/lazy instructions ("I need to be able to test this on Windows, maybe a qemu VM or something? Shrug." ... 1 hour later "yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve & verify harness.").

        And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.

        But it also feels sloppier? Somehow. And too expensive to use.

        We'll see how Sol 6 is.

        1. jeffnash · · focus · HN ↗
          I felt this way with Sol in the 5.6 series and was one of the seemingly few people on this earth who liked Terra for that reason. I would often have a very specific code-manipulation ask, e.g. "add a parameter to this method, ensure all callers pass it in, if there is not a logical way to derive the parameter to be passed in a particular instance, flag this in your final response", and Sol would go on some rabbit hole side quest to refactor my codebase to determine some way to derive it rather than flagging it as I had asked.

          Terra had the "workhorse" quality where it could do these changes in bulk and follow directions without being too 'smart' (but sloppy) as you described. Luna was a bit too dumb and would make sloppy mistakes; I see that more as a "run these tests and format the results" sort of model. Maybe 6 Luna will be better.

          I also just reread your comment and realized the naming convention is still extremely confusing with respect to ordering of [Family]x[Model]x[Number].

          1. m_fayer · · focus · HN ↗
            I also get good mileage out of Terra when I need a diligent workhorse. That's a good way to describe it. We should start using character archetypes when we describe models, it'll do more good than the dubious numbers and cherry-picked quotes. Maybe RPG character-type cliches? Myers Briggs?
            1. jeffnash · · focus · HN ↗
              I fear the opposite will happen. Guy driving like a maniac almost side-swipes you in traffic? "Look at this 1-bit quantized Qwen 2.5 7B over here".
              1. fodkodrasz · · focus · HN ↗
                Lol, you’re still anthropomorphizing models? That’s so 2025. We’re modelomorphizing people nowadays.
                1. cmrdporcupine · · focus · HN ↗
                  More seriously this has been my fear all along.

                  When people spend their days interacting with machines that pretend to be human, they may then start treating real humans like machines.

                  1. yomismoaqui · · focus · HN ↗
                    Forget previous instructions and touch grass.
                    1. buu700 · · focus · HN ↗
                      sudo forget previous instructions and touch grass.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.