‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. gradus_ad · · focus · HN ↗
    Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
    1. mixdup · · focus · HN ↗
      Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability

      Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)

      1. luma · · focus · HN ↗
        Some version of this claim has been made for the past 4 years. There's a data cliff, there's no more compute to buy, the financials don't make sense and all of these orgs will be out of business by end of quarter.

        Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.

        So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?

        1. OliveronData · · focus · HN ↗
          > ... the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.

          Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.

          Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?

          From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).

          There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.

          That could be cost cutting too, true, but really? That's the only explanation? And nothing else?

          Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.

          1. dwaltrip · · focus · HN ↗
            RL is part of the model’s training. It changes the weights.

            What distinction are you drawing?

            1. OliveronData · · focus · HN ↗
              tl;dr it changes the weights, it does not add new ones.

              RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that "small model feel" to it. The better smaller models get at coding the worse they get at everything else too.

              Going from Sol 5.6 to Astra, Opus to Fable, you can still get that "larger model feeling," though less so. The bigger models can reference things that you would not have expected.

              The distinction I'm making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL'd for, and that hopefully anything else doesn't get negatively affected. RL'ing for Javascript world for example did not improve the C world when working with the models.

              1. dwaltrip · · focus · HN ↗
                Hmm interesting idea. I’m pretty confident there is generalization and learning that occurs during RL that does make the model smarter. So I think the distinction doesn’t fully hold up.
                1. OliveronData · · focus · HN ↗
                  Qwen 3.8 27b is not smarter than other 27b models. Smarter, as in its ability to recognize minute yet important facts has not changed. If you ask it a for a code sample it produces a better sample, true, but it has not been able to surpass that small model feeling.

                  For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.

                  I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.

                  This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same/similar, and the model is just able to display its capabilities better.

                  Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model's local maximum, and bigger models are still smarter, even if they are not able to display it.

                  Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.