‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. gradus_ad · · focus · HN ↗
    Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic's rationale for IPOing this year.
    1. mixdup · · focus · HN ↗
      Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability

      Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)

      1. luma · · focus · HN ↗
        Some version of this claim has been made for the past 4 years. There's a data cliff, there's no more compute to buy, the financials don't make sense and all of these orgs will be out of business by end of quarter.

        Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.

        So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?

        1. dgellow · · focus · HN ↗
          Those points were true at the time and most are still true now. But they aren’t predictions.

          - it’s correct there isn’t much fresh data anymore

          - it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive

          - it’s correct the finances don’t make sense

          But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)

          1. moosehater · · focus · HN ↗
            I was thinking the same thing in terms of running out of data a few months ago. But aren't most gains in the past year+ due to reinforcement learning in some form? Which doesn't need "fresh data" per se, as the model effectively creates the data as it goes. As long as engineers can come up with proper environments, tasks/goals, rewards, and actions, I don't really see data being a limit to model improvement in an agentic sense. Maybe as a knowledge base
          2. dumberquestions · · focus · HN ↗

            [dead]

          3. JacobAsmuth · · focus · HN ↗
            The new hardware (TPU v8 and VR) are more expensive but they are significantly cheaper per flop. e.g. many multiples more performance for only 2x the price.

            If I have some ML workload to run I can buy $x of Blackwell chips or I can buy significantly less $ worth of Vera Rubin chips to get the same performance. That's the key thing to keep in mind when you're talking about financials.

          4. agoodusername63 · · focus · HN ↗
            The amount of irrationality I see in the economy with AI makes me more convinced that the wall street bankrollers know very well they're throwing money into a pit, but it's a pit they're gambling will turn into some world hunger ending AI (that will somehow also keep them making money off of scarcity)

            never mind that theres no guarantee we'll get that mythical AI. Never mind that the societal reformations would also impact their revenue numbers.

            1. dgellow · · focus · HN ↗
              I’m pretty convinced Wall Street is clueless and relies mostly on vibes. It’s the same people who thought all SaaS would become unnecessary after seeing a Claude code demo. Though eventually they will have to ask for the ROI, and that’s where the whole thing will calm down
        2. john_strinlai · · focus · HN ↗
          >Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.

          do you think it will be exponential forever?

          1. RobCat27 · · focus · HN ↗
            I think we'll eventually hit an information theoretic type of wall with physical hardware and GPUs and need a similar AI breakthrough as well as the development refinement of logical/physical qubits in the quantum computing space with some analogue to the transformer architecture to continue accelerating. However, I think there must be many years of development and refinement that can take place before that paradigm shift to overcome the physical compute wall is necessary. This is just my theory, but I'm young enough that I'm expecting with the rate that we are advancing, I will see AI / LLM analogues developed and run on a quantum computer in my lifetime.
            1. spathi_fwiffo · · focus · HN ↗
              I think the bottleneck will be the current one.

              Fabs.

              Either needing more fabs, new types of fabs, retooling existing fabs.

              All of that takes years.

              maybe we can design our way out of that too. But, I suppose that would be the similar breakthrough you are mentioning.

            2. [deleted] · · focus · HN ↗

              [deleted]

          2. sebzim4500 · · focus · HN ↗
            Eventually the heat death of the universe will come, so clearly any prediction made needs some kind of time frame attached.

            I think it's fully possible that it continues being exponential for decades like Moore's law did (and still is depending on exactly what you measure)

        3. trentnix · · focus · HN ↗
          Yep. I've made the claim (and been wrong). I was convinced the data cliff was going to be a real problem. Now I feel like we are on the cusp of having Tony Stark's Jarvis at our fingertips.

          What a time to be alive.

          1. neta1337 · · focus · HN ↗
            Incredible how many times I read similar comments over the years, containing 'on the cusp' and 'what a time to be alive'. Indeed, what a time - not a single user-facing thing on the internet has improved since then, considering the power tool we got. The most used web services get drowned in generated stuff and so are the users
            1. trentnix · · focus · HN ↗
              Not a single thing? In my house, we are using LLMs to:

              - plan youth soccer practices

              - develop well-formatted soccer game substitution schedules

              - build and ship software in languages I haven't used in 25 years on platforms I've never programmed for

              - do meal planning and build shopping lists

              - prepare grocery shopping carts

              - solicit medical advice

              - perform Garmin watch data analysis

              - administer devices (with SSH access) using natural language

              - avoid counterfeit soccer jersey purchases

              - create "Warrior Cat" graphic novels

              - make cartoon strips

              - troubleshoot appliances

              - manage finances

              - review accounting ledgers

              - diagnose malware infections

              - so much more

              And we do it all from a simple prompt that we can talk to if we choose.

              I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming.

              I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved.

              1. FiberBundle · · focus · HN ↗
                > I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming

                I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been.

                1. senderista · · focus · HN ↗
                  You can use LLMs to make your software better, but it's easier to use them to make it worse.
                2. semiquaver · · focus · HN ↗
                  Companies need to radically change to be able to take advantage. Most companies are afraid to do that and are letting their engineers serve as slow meat proxies, doing software development basically the same way as before.
                3. robryan · · focus · HN ↗
                  Probably because it is in flux, there is a large scale reorganisation of software around agents going on.
        4. digdugdirk · · focus · HN ↗
          The difference now is that they've hit the "good enough" point. LLMs are a tool, and that tool is useful but not incredibly valuable unto itself.

          To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.

          Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.

          So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.

          1. willchis · · focus · HN ↗
            This is how I feel about it. I've stopped looking at all the scores of new releases and just look at the price to see how much usage I can get in a month. Seems like I'm not the only one either, from comments above like

            > "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_

          2. famouswaffles · · focus · HN ↗
            >The difference now is that they've hit the "good enough" point.

            In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D asset creation are all things Astra is >> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong.

            1. digdugdirk · · focus · HN ↗
              Right, that's exactly my point though. Those are absolutely valuable use cases. But are they useful enough to justify a trillion dollar valuation? Or is most of the economic value in the stuff that already exists, and can be performed nearly as well by qwen/deepseek/kimi/GLM/etc?

              And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?

              1. famouswaffles · · focus · HN ↗
                >Right, that's exactly my point though. Those are absolutely valuable use cases. But are they useful enough to justify a trillion dollar valuation?

                Replacing white collar work would be worth dozens of trillions of dollars. Software is not the only valuable job that can be done on a computer.

                OpenAI and Anthropic already have what it takes right now to become trillion dollar companies even if the above doesn't materialize.

                Chatgpt is used by a billion people every week. Their ads program hit $1B ARR in 200 days. And Anthropic is growing so fast they're in track to hit an Annual Revenue Run Rate of $100B before they IPO.

                >And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?

                How long until...you could say that about the capabilities of past models but OpenAI still dwarf everyone else in consumer usage, and Anthropic still have enterprise usage on lock. In the end, neither the billion+ users of gpt or the enterprise customers are going to give a shit. And specialized models often perform worse than generalized ones.

        5. OliveronData · · focus · HN ↗
          > ... the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.

          Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.

          Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?

          From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).

          There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.

          That could be cost cutting too, true, but really? That's the only explanation? And nothing else?

          Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.

          1. luma · · focus · HN ↗
            I didn't use the word LLM. I'm talking AI capability, you're focused on this or that current approach to AI. I think it's fair to assume that the approach will change as new ideas are learned, new and more hardware will be purchased and applied to the problem, and then capabilities will (for now) continue on their exponential curve, same as it has gone for the past several years.

            These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.

            1. neta1337 · · focus · HN ↗
              It is a bit harsh to call it knocking down considering all facts
          2. dwaltrip · · focus · HN ↗
            RL is part of the model’s training. It changes the weights.

            What distinction are you drawing?

            1. OliveronData · · focus · HN ↗
              tl;dr it changes the weights, it does not add new ones.

              RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that "small model feel" to it. The better smaller models get at coding the worse they get at everything else too.

              Going from Sol 5.6 to Astra, Opus to Fable, you can still get that "larger model feeling," though less so. The bigger models can reference things that you would not have expected.

              The distinction I'm making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL'd for, and that hopefully anything else doesn't get negatively affected. RL'ing for Javascript world for example did not improve the C world when working with the models.

              1. dwaltrip · · focus · HN ↗
                Hmm interesting idea. I’m pretty confident there is generalization and learning that occurs during RL that does make the model smarter. So I think the distinction doesn’t fully hold up.
                1. OliveronData · · focus · HN ↗
                  Qwen 3.8 27b is not smarter than other 27b models. Smarter, as in its ability to recognize minute yet important facts has not changed. If you ask it a for a code sample it produces a better sample, true, but it has not been able to surpass that small model feeling.

                  For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.

                  I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.

                  This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same/similar, and the model is just able to display its capabilities better.

                  Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model's local maximum, and bigger models are still smarter, even if they are not able to display it.

                  Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?

            2. kqr · · focus · HN ↗
              I think the distinction is between "improving the g factor" and "adapting a given level of g to perform certain types of work better".
        6. interestpiqued · · focus · HN ↗
          4 years is not that long in the grand scheme of things to be fair
        7. dcchambers · · focus · HN ↗

          [dead]

        8. chamomeal · · focus · HN ↗
          Has it been exponential this whole time? I feel like GPT-4 was pretty dang good. Maybe it’s rose tinted glasses cause I could finally have a bot write my dockerfiles and bash scripts, which knocked my socks off
        9. holbrad · · focus · HN ↗
          I think this is just a case of the bitter lesson that increasing compute just makes all these predictions meaningless. LLMs just keep going when everyone predicts them to fail constantly.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.