‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1065 points · 952 comments · crorella

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. hlynurd · · focus · HN ↗
    Weren't there headlines just yesterday that they weren't releasing this due to safety concerns?
    1. tedsanders · · focus · HN ↗
      GPT-6.1 Astra is what those headlines referred to. This is GPT-6.1 Sol.
      1. thejazzman · · focus · HN ↗
        "no homerS -- we're allowed to have one"

        <a href="https:&#x2F;&#x2F;amphetamem.es&#x2F;meme?id=the-simpsons_06_12_71&amp;text=We%27re+allowed+to+have+one.&amp;j=" rel="nofollow">https:&#x2F;&#x2F;amphetamem.es&#x2F;meme?id=the-simpsons_06_12_71&amp;text=We%...

      2. afroboy · · focus · HN ↗
        I lost track of those naming and numbers, AI is moving so fast.
    2. pkulak · · focus · HN ↗
      That was 6.1 Astra. And I&#x27;m assuming it&#x27;s being tabled because it still doesn&#x27;t beat match Opus 5.5.

      This is a decent win though, if it really is better. 6-sol was really no good, at least in my work.

      1. cmrdporcupine · · focus · HN ↗
        6 Sol was worse than 5.6 Sol from my own experiences. Far worse.

        Will see if this remedies things.

        1. pkulak · · focus · HN ↗
          Yup, same experience here. I used it for one day, spent the next day fixing its lousy code, then went back to 5.6.
  2. prodigycorp · · focus · HN ↗
    API price cuts were obvious once they made their announcement changing how usage is counted.

    These moves all make sense when you take into account the enterprise market.

    <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49889873">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49889873

  3. aaronbrethorst · · focus · HN ↗
    Shots fired, half the price of Opus 5.5.
  4. t-sauer · · focus · HN ↗
    Wasn&#x27;t 6 released like last week? I can&#x27;t keep up anymore.
    1. nsingh2 · · focus · HN ↗
      Yes but it was underwhelming, so the seem to have rushed 6.1 Sol out.
    2. SirMaster · · focus · HN ↗
      Do you need to? Do you always keep up with all the version bumps on the software you use?
    3. algoth1 · · focus · HN ↗
      It was so underwhelming that it didn&#x27;t even make it to chatgpt chat interface
      1. oh_no · · focus · HN ↗
        Sol 6 is in there? You may be on Enterprise where it didn&#x27;t roll out by default and comes out in a week or so. (Which is a weird and bad change to their model releases.)
        1. RugnirViking · · focus · HN ↗
          we have it in our enterprise. The admins have to individually approve new model releases. Ours don&#x27;t enable astra :(
  5. gradus_ad · · focus · HN ↗
    Ominous for the industry and investors that token price is becoming the main battleground. Could be Anthropic&#x27;s rationale for IPOing this year.
    1. nojito · · focus · HN ↗
      Great for the consumer.

      I remember when bandwidth was super expensive and now it’s dirt cheap.

      1. vanviegen · · focus · HN ↗
        Not an AWS customer, I take it? :-)
      2. iAMkenough · · focus · HN ↗
        That&#x27;s relative to where you live.

        Consumers are now saying the new pricing is not so great. <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49897236">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49897236

    2. djfjkfkffkkf · · focus · HN ↗
      China will do to llms what they did to german cars
      1. Razengan · · focus · HN ↗
        Why make a new account just to post this comment?

        It&#x27;s not even anything controversial..

        1. simlevesque · · focus · HN ↗
          They may work for one of the big AI labs.
          1. necovek · · focus · HN ↗
            Or a German car company?
            1. Razengan · · focus · HN ↗
              Maybe they&#x27;re a Chinese AI car?
        2. neta1337 · · focus · HN ↗
          Are there no new users expected?
          1. alch- · · focus · HN ↗
            Called djfjkfkffkkf? Not really.
      2. bogrollben · · focus · HN ↗
        I guess I&#x27;m out of touch. What did china do to german cars?
        1. Razengan · · focus · HN ↗
          What they&#x27;ll do to LLMs. Keep up
          1. coderenegade · · focus · HN ↗
            There&#x27;s still a sizable gap between OpenAI, Anthropic, and the Chinese labs. If anything, it&#x27;s getting bigger. We haven&#x27;t seen Chinese models solve the types of problems that ChatGPT and Claude are able to solve.
            1. bendtb · · focus · HN ↗
              But 95% of all problems are fairly banal, i.e. if the AI in my dish washer, fridge, radio, bicycle computer, etc. just have GPT 5.5 intelligence for making sure temperature, water, direction, etc. are 20% better controlled than yesterday it will be an enormous win for ordinary people (and the ressources we consume).
              1. 2yrrr · · focus · HN ↗
                Correct the implicit gamble of frontier labs is to displace humans - and the only firm you should trust with this is an American one.

                It’s not happening and the imminent bust is coming. Strap in while the music gets turned up (dots, ipo etc) and people decide to leave the partayyy!

        2. CamperBob2 · · focus · HN ↗
          Outcompeted them badly. Sent Porsche packing and BMW bawling.
          1. zigzag312 · · focus · HN ↗
            By being subsidised by the government.
            1. tngranados · · focus · HN ↗
              [delayed]
              1. zigzag312 · · focus · HN ↗
                Yes, but not nearly to the same extent.

                First, the subsidies to consumers for electric vehicles in Germany apply to all cars, not just those built in Europe. This effectively subsidizes the competition from China.

                As I didn&#x27;t quickly find any source making the direct comparison, I asked LLM to research it (I apologize for for this, but doing it manually would take too much time).

                   BYD: ~15% producer-focused economic benefit under the Commission&#x27;s methodology; 17.0% including the legacy NEV fiscal scheme.
                
                   VW Europe: ~0.2–0.7% is my best public-data-based estimate of currently observable and allocatable producer support; roughly 0.7–1.3% if we deliberately make aggressive assumptions favorable to VW.
                
                   VW extreme stress test: ~2.5–3.2%, obtained by implausibly allocating essentially all VW Group grants and tax credits to European BEVs.
                
                EU Commission&#x27;s investigation calculated countervailable subsidy rates for Chinese BEVs: see &quot;3.10.3. Calculation of subsidy rates&quot; for a aggregate subsidy rates.

                <a href="https:&#x2F;&#x2F;eur-lex.europa.eu&#x2F;legal-content&#x2F;EN&#x2F;TXT&#x2F;?qid=1738249752833&amp;uri=CELEX%3A32024R2754" rel="nofollow">https:&#x2F;&#x2F;eur-lex.europa.eu&#x2F;legal-content&#x2F;EN&#x2F;TXT&#x2F;?qid=17382497...

            2. CamperBob2 · · focus · HN ↗
              That argument doesn&#x27;t hold much water in my opinion. VW AG is partially state-owned, with a double-digit percentage held by the federal state of Lower Saxony.

              Meanwhile, in the US, companies like Boeing, GM, and Intel will never be allowed to experience more than minor financial inconvenience before the government bails them out with protectionism, loans, and outright subsidies.

              I just don&#x27;t see a material difference between how the Chinese government treats their strategically-important industries and the way we do here in the West.

              1. zigzag312 · · focus · HN ↗
                &gt; VW AG is partially state-owned, with a double-digit percentage held by the federal state of Lower Saxony.

                Yes, but those are two separate things. Lower Saxony owning part of VW doesn’t mean VW gets extra public money because of it.

                High subsidy rates alter market dynamics. Can we agree on that?

                1. CamperBob2 · · focus · HN ↗
                  For sure.
                  1. zigzag312 · · focus · HN ↗
                    The biggest subsidies for EV in Germany are subsidies to consumers, which apply also to Chinese cars.

                    Subsidies to manufacturers in Germany are much smaller than both: subsidies to consumers in Germany, and subsidies to manufacturers in China.

                    So, there are different subsidy types with different subsidy rates which all impact the market. Higher rates usually have a higher impact, but type of subsidy can also change what kind of an impact they have. Any comparison quickly becomes complex.

                    I posted some numbers in an other post [0], but the numbers are not exact.

                    Subsidy is also one of many factors. Of course you also need capable people and manufacturing capabilities which China also has. Subsidy alone is not enough, but if needed capabilities exist, subsidy at high rates can help tip the scales.

                    [0] <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49920502">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49920502

      3. ok123456 · · focus · HN ↗
        We can only hope.
    3. mixdup · · focus · HN ↗
      Another piece of evidence on the pile that the sudden panic and desire to &quot;slow down&quot; is because they&#x27;re hitting the plateau on capability

      Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)

      1. semiquaver · · focus · HN ↗
        What universe do you live in that you can look at the past six months and see anything like a plateau in capability?

        This is brain-rotted zitron-conspiracy territory, utterly at odds with reality.

        1. ActionHank · · focus · HN ↗
          Have we honestly seen that great a leap in the last 6 months, or just better application of what we had 6 months before that.

          We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it&#x27;s more of the same.

          1. CuriouslyC · · focus · HN ↗
            The difference between 6 months ago frontier and now frontier in 3d modelling, graphics and video editing is night and day.
            1. ActionHank · · focus · HN ↗
              Just because there are new capabilities, doesn&#x27;t mean they&#x27;ve pushed passed the plateau, they&#x27;ve just expanded where the previous solutions work.

              We&#x27;ve gone from 80% in some places to 80% in some more places.

              1. CuriouslyC · · focus · HN ↗
                &gt; We&#x27;ve gone from 80% in some places to 80% in some more places.

                Any area that is verifiable will trend inexorably towards 100% over time. In unverifiable areas, it&#x27;ll always be &quot;80%&quot; because the ubiquity of &quot;AI&quot; style erodes its value, and &quot;&gt;80%&quot; for unverifiable things involves fashion, cachet and &quot;vibes&quot; that humans will probably never knowingly let it have.

                1. ActionHank · · focus · HN ↗
                  Just checked your website, you really drank all the koolaid huh?
              2. coderenegade · · focus · HN ↗
                I don&#x27;t think the math community thinks there&#x27;s been a plateau. The models have gone from being a joke to being able to crank out proofs to research grade problems.
          2. usef- · · focus · HN ↗
            Statements like these remind me how much each of us are in our own unique information bubbles. People seemed very excited about analysing the differences of each model on my feeds.

            Have you used recent ones on any large projects or coding issues? They&#x27;ve improved tremendously lately.

        2. mixdup · · focus · HN ↗
          Not that they&#x27;ve hit it but that they are approaching it. The time to panic and steer the narrative is before you hit the iceberg, not after
          1. helloplanets · · focus · HN ↗
            The whole &quot;pace the frontier&quot; thing originated from Anthropic.

            With Opus 5.5, it doesn&#x27;t seem like model improvement is plateuaing. And Fable 5.5 will likely be dropped this or next week.

            You can only imagine what they&#x27;ve got going internally.

        3. phoghed · · focus · HN ↗
          People have been saying this since GPT-4.
        4. arctic-true · · focus · HN ↗
          Most of the impressive accomplishments we’ve seen in the last few months have been the result of huge agent swarms working together and brute-forcing solutions, not massive leaps in intelligence from standalone models. That is still an improvement in the usefulness and power of the technology, but it is NOT evidence that model intelligence is increasing faster than before.
          1. famouswaffles · · focus · HN ↗
            I don&#x27;t have any access to any agent swarms (and neither do most) and i still think the models have obviously improved massively in standalone intelligence. Of course they have, agent swarms are not magic. You can swarm all you want around GPT-4 era models and you&#x27;ll get nowhere. And i&#x27;ve never seen the term &#x27;brute-force&#x27; more abused than these LLM discussions. Basically none of the results have been brute force.
            1. semiquaver · · focus · HN ↗
              Agreed. You can’t “brute force” reality, which has an infinitely large state space. A million monkeys won’t write Shakespeare and all that
          2. CamperBob2 · · focus · HN ↗
            &quot;This machine-intelligence stuff is overrated, they are just using &lt;insert particular machine-intelligence technique here&gt;&quot; isn&#x27;t the resounding verdict it may have sounded like when you typed it.
          3. usef- · · focus · HN ↗
            Do you use them? I feel like you&#x27;re talking about the news headlines. But in daily, direct, individual usage (one model at a time) I&#x27;ve found they&#x27;ve improved tremendously on the last few months.
            1. arctic-true · · focus · HN ↗
              Yes, I do. They’ve definitely gotten better, I wouldn’t dispute that. But we haven’t gone from useless toys to Skynet in the last 6 months, progress is slower and steadier than that.
      2. serf · · focus · HN ↗
        &gt;Another piece of evidence on the pile that the sudden panic and desire to &quot;slow down&quot; is because they&#x27;re hitting the plateau on capability

        if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.

        We&#x27;re still improving transistors on a somewhat routine basis.

        1. mixdup · · focus · HN ↗
          The plateau doesn&#x27;t have to be perfectly flat, but it&#x27;s not a straight line upward anymore either (kind of like our work on transistors, where we&#x27;ve kind of hit the bounds of speed in clock cycles but are improving on miniaturization and power efficiency)
          1. password54321 · · focus · HN ↗
            It took 5 years to solve ARC-AGI 1, 1 year to solve ARC-AGI 2 and 6 months to solve ARC-AGI 3.
            1. delillos · · focus · HN ↗
              Those version numbers don&#x27;t necessarily correspond to equal increases in &quot;difficulty&quot;, though.
              1. password54321 · · focus · HN ↗
                Correct, the benchmark became exponentially more difficult.
              2. JacobAsmuth · · focus · HN ↗
                Very true, the sharp increase in difficulty (as measured by human passrate plummeting from 1-&gt;2 and again from 2-&gt;3) gives an even more stark view of AI capabilities over time.
      3. LPisGood · · focus · HN ↗
        Nvidia can start putting weights in silicon if model development slows down.
      4. theturtletalks · · focus · HN ↗
        I think they are hitting compute restrictions. And buying compute right now can be 3-4X. And the costs are increasing. If they train a larger model and demand is high, that’s a lot of compute for Codex subscriptions, which is a loss leader for them. Especially Pro 20X which they just nerfed to 10X.
      5. colechristensen · · focus · HN ↗
        &gt;Another piece of evidence on the pile that the sudden panic and desire to &quot;slow down&quot; is because they&#x27;re hitting the plateau on capability

        I think it&#x27;s more a token-cost-demand plateau. They&#x27;ve reached the scale and investor trillions to which they can&#x27;t 10x the hardware cost of inference any more. They can&#x27;t afford to compete by eating costs and there isn&#x27;t appetite for more expensive inference.

        So in order that they don&#x27;t bankrupt each other they&#x27;re looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.

        There&#x27;s a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.

        Maybe it&#x27;s good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.

      6. CuriouslyC · · focus · HN ↗
        It&#x27;s not so much that they&#x27;re hitting a plateau in capability, as we&#x27;re saturating long horizon benchmarks and it&#x27;s not greatly improving general usability. On the other hand, newer models have been amazing for people interested in 3d, graphics, video editing, etc. The difference between Opus 5.5&#x2F;Astra and earlier models is night and day even if for many coding tasks they&#x27;re not a revolution.
        1. omalled · · focus · HN ↗
          I agree that they&#x27;re not hitting a plateau and I see it in my reserach. I had a math&#x2F;code benchmark paper [1] at NeurIPS last year that is still unsaturated. At the time of writing the paper, the best model was o3, which was scoring 3-4%. By the time NeurIPS came around, GPT-5.2 was the latest model but it was getting similar scores to o3. The models were still in the flat part of the usual hockey stick curve. The newer models are getting into the steep part. I evaluated gpt-5.6-sol+codex a week or two ago and it got ~16%. Astra+codex got ~24%.

          On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered &quot;hard&quot; [2]. It produced a lean proof that the formula is correct, but I&#x27;m just starting to learn lean and don&#x27;t have enough expertise to check it.

          [1] <a href="https:&#x2F;&#x2F;proceedings.neurips.cc&#x2F;paper_files&#x2F;paper&#x2F;2025&#x2F;hash&#x2F;c7c2c3ab977316105f9c3b3fdf769f0e-Abstract-Datasets_and_Benchmarks_Track.html" rel="nofollow">https:&#x2F;&#x2F;proceedings.neurips.cc&#x2F;paper_files&#x2F;paper&#x2F;2025&#x2F;hash&#x2F;c... [2] <a href="https:&#x2F;&#x2F;oeis.org&#x2F;A000530" rel="nofollow">https:&#x2F;&#x2F;oeis.org&#x2F;A000530

      7. xienze · · focus · HN ↗
        &gt; sudden panic and desire to &quot;slow down&quot; is because they&#x27;re hitting the plateau on capability

        I don&#x27;t think that&#x27;s the motivation, it&#x27;s because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman&#x27;s agreement (they won&#x27;t), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they&#x27;ve gotten themselves into with the emphasis on being best, with premium prices to match.

        1. redanddead · · focus · HN ↗
          The game is already over
      8. azan_ · · focus · HN ↗
        &gt; Another piece of evidence on the pile that the sudden panic and desire to &quot;slow down&quot; is because they&#x27;re hitting the plateau on capability

        People were talking about plateau for years already.

      9. sebzim4500 · · focus · HN ↗
        Is there anything that could happen that you wouldn&#x27;t use as evidence that they are hitting a plateau?

        It just seems like these claims are constant and looking back the calls of &#x27;plateau&#x27; between 2023 and 2025 were clearly false, why should we think it&#x27;s different now?

        1. djdjdkdkfk · · focus · HN ↗
          I am a lawyer not a coder, but for me the new models make the exact dumb mistakes they did in 2023. Everytime I come here I feel I&#x27;m in an alternate reality.
          1. RussianCow · · focus · HN ↗
            [delayed]
          2. robryan · · focus · HN ↗
            You get no better result from Opus 5.5 than Opus 4.1?
          3. epihelix · · focus · HN ↗
            IANAL, but it sounds as though law hallucinations are the final frontier. Sorry.

            But coding-wise, models keep getting better and cheaper. You can train for code correctness in a way you can&#x27;t train for legal correctness, and you can test your code in an agentic loop in a way you can&#x27;t test a legal opinion.

            Hence your alternative reality.

            (All that said, 2023 was GPT-4 territory. GPT-4o wasn&#x27;t released until 2024. No matter what question you&#x27;re asking, I struggle to believe you wouldn&#x27;t notice the difference between GPT-4 and the current frontier model set. You can download and run any number of sub-27B local models that will be better than GPT-4. The pace of change in this field really has been insane.)

          4. rspeele · · focus · HN ↗
            Both the rate of improvement and the actual effectiveness of LLMs is much higher in coding than in other fields. An LLM+coding harness can try something wrong and get rapidly, harmlessly caught by a compiler or by a unit test. Then it usually gets it right on the 2nd or 3rd attempt. This all happens before the user who requested the change is told it&#x27;s done. And a lot of it goes right back into the training dataset for the next model release.

            In law if a model makes a mistake it pretty much takes a human to catch it, which is a vastly slower, riskier, and more frustrating feedback loop. So it doesn&#x27;t surprise me they can stay dumb in that field while getting massively more capable in coding with multiple &quot;step change&quot; releases in the past 365 days.

      10. luma · · focus · HN ↗
        Some version of this claim has been made for the past 4 years. There&#x27;s a data cliff, there&#x27;s no more compute to buy, the financials don&#x27;t make sense and all of these orgs will be out of business by end of quarter.

        Not once has any of these predictions come true, the pace of progress has continued on it&#x27;s exponential trajectory since ChatGPT first came to the public&#x27;s attention.

        So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?

        1. dgellow · · focus · HN ↗
          Those points were true at the time and most are still true now. But they aren’t predictions.

          - it’s correct there isn’t much fresh data anymore

          - it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive

          - it’s correct the finances don’t make sense

          But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)

          1. moosehater · · focus · HN ↗
            I was thinking the same thing in terms of running out of data a few months ago. But aren&#x27;t most gains in the past year+ due to reinforcement learning in some form? Which doesn&#x27;t need &quot;fresh data&quot; per se, as the model effectively creates the data as it goes. As long as engineers can come up with proper environments, tasks&#x2F;goals, rewards, and actions, I don&#x27;t really see data being a limit to model improvement in an agentic sense. Maybe as a knowledge base
          2. dumberquestions · · focus · HN ↗

            [dead]

          3. JacobAsmuth · · focus · HN ↗
            The new hardware (TPU v8 and VR) are more expensive but they are significantly cheaper per flop. e.g. many multiples more performance for only 2x the price.

            If I have some ML workload to run I can buy $x of Blackwell chips or I can buy significantly less $ worth of Vera Rubin chips to get the same performance. That&#x27;s the key thing to keep in mind when you&#x27;re talking about financials.

          4. agoodusername63 · · focus · HN ↗
            The amount of irrationality I see in the economy with AI makes me more convinced that the wall street bankrollers know very well they&#x27;re throwing money into a pit, but it&#x27;s a pit they&#x27;re gambling will turn into some world hunger ending AI (that will somehow also keep them making money off of scarcity)

            never mind that theres no guarantee we&#x27;ll get that mythical AI. Never mind that the societal reformations would also impact their revenue numbers.

            1. dgellow · · focus · HN ↗
              I’m pretty convinced Wall Street is clueless and relies mostly on vibes. It’s the same people who thought all SaaS would become unnecessary after seeing a Claude code demo. Though eventually they will have to ask for the ROI, and that’s where the whole thing will calm down
        2. john_strinlai · · focus · HN ↗
          &gt;Not once has any of these predictions come true, the pace of progress has continued on it&#x27;s exponential trajectory since ChatGPT first came to the public&#x27;s attention.

          do you think it will be exponential forever?

          1. RobCat27 · · focus · HN ↗
            I think we&#x27;ll eventually hit an information theoretic type of wall with physical hardware and GPUs and need a similar AI breakthrough as well as the development refinement of logical&#x2F;physical qubits in the quantum computing space with some analogue to the transformer architecture to continue accelerating. However, I think there must be many years of development and refinement that can take place before that paradigm shift to overcome the physical compute wall is necessary. This is just my theory, but I&#x27;m young enough that I&#x27;m expecting with the rate that we are advancing, I will see AI &#x2F; LLM analogues developed and run on a quantum computer in my lifetime.
            1. spathi_fwiffo · · focus · HN ↗
              I think the bottleneck will be the current one.

              Fabs.

              Either needing more fabs, new types of fabs, retooling existing fabs.

              All of that takes years.

              maybe we can design our way out of that too. But, I suppose that would be the similar breakthrough you are mentioning.

            2. [deleted] · · focus · HN ↗

              [deleted]

          2. sebzim4500 · · focus · HN ↗
            Eventually the heat death of the universe will come, so clearly any prediction made needs some kind of time frame attached.

            I think it&#x27;s fully possible that it continues being exponential for decades like Moore&#x27;s law did (and still is depending on exactly what you measure)

        3. trentnix · · focus · HN ↗
          Yep. I&#x27;ve made the claim (and been wrong). I was convinced the data cliff was going to be a real problem. Now I feel like we are on the cusp of having Tony Stark&#x27;s Jarvis at our fingertips.

          What a time to be alive.

          1. neta1337 · · focus · HN ↗
            Incredible how many times I read similar comments over the years, containing &#x27;on the cusp&#x27; and &#x27;what a time to be alive&#x27;. Indeed, what a time - not a single user-facing thing on the internet has improved since then, considering the power tool we got. The most used web services get drowned in generated stuff and so are the users
            1. trentnix · · focus · HN ↗
              Not a single thing? In my house, we are using LLMs to:

              - plan youth soccer practices

              - develop well-formatted soccer game substitution schedules

              - build and ship software in languages I haven&#x27;t used in 25 years on platforms I&#x27;ve never programmed for

              - do meal planning and build shopping lists

              - prepare grocery shopping carts

              - solicit medical advice

              - perform Garmin watch data analysis

              - administer devices (with SSH access) using natural language

              - avoid counterfeit soccer jersey purchases

              - create &quot;Warrior Cat&quot; graphic novels

              - make cartoon strips

              - troubleshoot appliances

              - manage finances

              - review accounting ledgers

              - diagnose malware infections

              - so much more

              And we do it all from a simple prompt that we can talk to if we choose.

              I&#x27;ve built more (and better) software in the past month than I did in any given year in the 30+ years I&#x27;ve been programming.

              I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there&#x27;s no good reason at all to be pessimistic about how quickly this has improved.

              1. FiberBundle · · focus · HN ↗
                &gt; I&#x27;ve built more (and better) software in the past month than I did in any given year in the 30+ years I&#x27;ve been programming

                I feel similarly, but I think it&#x27;s a valid question. Why is all the software I&#x27;m using not getting better? To be honest, I feel it&#x27;s more buggy than it&#x27;s ever been.

                1. senderista · · focus · HN ↗
                  You can use LLMs to make your software better, but it&#x27;s easier to use them to make it worse.
                2. semiquaver · · focus · HN ↗
                  Companies need to radically change to be able to take advantage. Most companies are afraid to do that and are letting their engineers serve as slow meat proxies, doing software development basically the same way as before.
                3. robryan · · focus · HN ↗
                  Probably because it is in flux, there is a large scale reorganisation of software around agents going on.
        4. digdugdirk · · focus · HN ↗
          The difference now is that they&#x27;ve hit the &quot;good enough&quot; point. LLMs are a tool, and that tool is useful but not incredibly valuable unto itself.

          To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we&#x27;ve gone from that to a 3-axis CNC mill. Now we&#x27;ve added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn&#x27;t want&#x2F;need them so they&#x27;re competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.

          Oh, and we&#x27;ve bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.

          So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.

          1. willchis · · focus · HN ↗
            This is how I feel about it. I&#x27;ve stopped looking at all the scores of new releases and just look at the price to see how much usage I can get in a month. Seems like I&#x27;m not the only one either, from comments above like

            &gt; &quot;Opus 5.5 is so good that I don&#x27;t want it to be replaced anytime soon. Stop training models[...]&quot;_

          2. famouswaffles · · focus · HN ↗
            &gt;The difference now is that they&#x27;ve hit the &quot;good enough&quot; point.

            In some aspects sure, but in others no. Open AI&#x27;s goal is to build &quot;highly autonomous systems that outperform humans at most economically valuable work.&quot; and Astra was a big jump in that. There still isn&#x27;t a better model for computer use and vision&#x2F;spatial work. Driving, Operating Robots, Video Editing, 3D asset creation are all things Astra is &gt;&gt; at than any other model. I&#x27;m sure you don&#x27;t care about any of that so it&#x27;s easy enough to slip by you but this analogy - &quot;Now we&#x27;ve added a 4th and 5th axis, which is great for the 2% of parts that need that functionality.&quot; is dead wrong.

            1. digdugdirk · · focus · HN ↗
              Right, that&#x27;s exactly my point though. Those are absolutely valuable use cases. But are they useful enough to justify a trillion dollar valuation? Or is most of the economic value in the stuff that already exists, and can be performed nearly as well by qwen&#x2F;deepseek&#x2F;kimi&#x2F;GLM&#x2F;etc?

              And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?

              1. famouswaffles · · focus · HN ↗
                &gt;Right, that&#x27;s exactly my point though. Those are absolutely valuable use cases. But are they useful enough to justify a trillion dollar valuation?

                Replacing white collar work would be worth dozens of trillions of dollars. Software is not the only valuable job that can be done on a computer.

                OpenAI and Anthropic already have what it takes right now to become trillion dollar companies even if the above doesn&#x27;t materialize.

                Chatgpt is used by a billion people every week. Their ads program hit $1B ARR in 200 days. And Anthropic is growing so fast they&#x27;re in track to hit an Annual Revenue Run Rate of $100B before they IPO.

                &gt;And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?

                How long until...you could say that about the capabilities of past models but OpenAI still dwarf everyone else in consumer usage, and Anthropic still have enterprise usage on lock. In the end, neither the billion+ users of gpt or the enterprise customers are going to give a shit. And specialized models often perform worse than generalized ones.

        5. OliveronData · · focus · HN ↗
          &gt; ... the pace of progress has continued on it&#x27;s exponential trajectory since ChatGPT first came to the public&#x27;s attention.

          Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn&#x27;t as attentive as Fable in long term writing for example.

          Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?

          From where I&#x27;m standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat&#x2F;Work split).

          There&#x27;s a race from OpenAI to serve dumber models on chat. I&#x27;m not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely &quot;chatted&quot; with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I&#x27;m guessing they will announce that later during the dev days.

          That could be cost cutting too, true, but really? That&#x27;s the only explanation? And nothing else?

          Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.

          1. luma · · focus · HN ↗
            I didn&#x27;t use the word LLM. I&#x27;m talking AI capability, you&#x27;re focused on this or that current approach to AI. I think it&#x27;s fair to assume that the approach will change as new ideas are learned, new and more hardware will be purchased and applied to the problem, and then capabilities will (for now) continue on their exponential curve, same as it has gone for the past several years.

            These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.

            1. neta1337 · · focus · HN ↗
              It is a bit harsh to call it knocking down considering all facts
          2. dwaltrip · · focus · HN ↗
            RL is part of the model’s training. It changes the weights.

            What distinction are you drawing?

            1. OliveronData · · focus · HN ↗
              tl;dr it changes the weights, it does not add new ones.

              RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that &quot;small model feel&quot; to it. The better smaller models get at coding the worse they get at everything else too.

              Going from Sol 5.6 to Astra, Opus to Fable, you can still get that &quot;larger model feeling,&quot; though less so. The bigger models can reference things that you would not have expected.

              The distinction I&#x27;m making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL&#x27;d for, and that hopefully anything else doesn&#x27;t get negatively affected. RL&#x27;ing for Javascript world for example did not improve the C world when working with the models.

              1. dwaltrip · · focus · HN ↗
                Hmm interesting idea. I’m pretty confident there is generalization and learning that occurs during RL that does make the model smarter. So I think the distinction doesn’t fully hold up.
                1. OliveronData · · focus · HN ↗
                  Qwen 3.8 27b is not smarter than other 27b models. Smarter, as in its ability to recognize minute yet important facts has not changed. If you ask it a for a code sample it produces a better sample, true, but it has not been able to surpass that small model feeling.

                  For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.

                  I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.

                  This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same&#x2F;similar, and the model is just able to display its capabilities better.

                  Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model&#x27;s local maximum, and bigger models are still smarter, even if they are not able to display it.

                  Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?

            2. kqr · · focus · HN ↗
              I think the distinction is between &quot;improving the g factor&quot; and &quot;adapting a given level of g to perform certain types of work better&quot;.
        6. interestpiqued · · focus · HN ↗
          4 years is not that long in the grand scheme of things to be fair
        7. dcchambers · · focus · HN ↗

          [dead]

        8. chamomeal · · focus · HN ↗
          Has it been exponential this whole time? I feel like GPT-4 was pretty dang good. Maybe it’s rose tinted glasses cause I could finally have a bot write my dockerfiles and bash scripts, which knocked my socks off
        9. holbrad · · focus · HN ↗
          I think this is just a case of the bitter lesson that increasing compute just makes all these predictions meaningless. LLMs just keep going when everyone predicts them to fail constantly.
      11. nbardy · · focus · HN ↗
        Have you even tried the new models? Opus 5.5 is a clear leap. Go back a year ago and try them and tell me there is any sort of plateau.
    4. jorblumesea · · focus · HN ↗
      This is literally the plan, open weight models are something like 60% of token spend, and it will get worse. many companies now have model gateways where you can slot in cheaper models via cli for cheaper. we&#x27;ve been using glm 5.x and it&#x27;s pretty close to SOTA frontier models.

      it&#x27;s also why there have been so many calls for regulation and slowdowns.

      1. LeBit · · focus · HN ↗
        Yup.

        I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.

        I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.

        1. nozzlegear · · focus · HN ↗
          [delayed]
      2. 0cf8612b2e1e · · focus · HN ↗
        There is already tooling to automatically pick models within an organization. Eventually it could be as easy as flipping a switch in group policy that forces everyone to switch to the cheaper models.

        Insane pricing pressure on the horizon.

    5. jimbob45 · · focus · HN ↗
      Pretty standard business to identify and compete on every axis (cost, speed, intelligence, etc). Often, nobody will be able to maximize every axis so you end up with a polyhedron derived from the axes where there’s a niche for everyone.

      DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.

    6. simianwords · · focus · HN ↗
      ?! this model launch was around 10% of the dev day and the other time was spent on Dots and things other than models.
      1. minimaxir · · focus · HN ↗
        That makes sense. There&#x27;s not really much else you can say about it.
    7. eli · · focus · HN ↗
      It would be weird if consumers were completely price insensitive.
  6. amelius · · focus · HN ↗
    If these models are so smart, can&#x27;t _they_ select the right model for each task?
    1. skulk · · focus · HN ↗
      the right model for the task is the one that transfers the maximum amount of USD from your pocket to the provider&#x27;s bank account.
      1. amelius · · focus · HN ↗
        No because then I&#x27;ll go to the competition.
    2. aleph_minus_one · · focus · HN ↗
      Why don&#x27;t you simply ask the respective model which model is best for a specific task? :-)
      1. amelius · · focus · HN ↗
        Because it is more work?
    3. mholm · · focus · HN ↗
      Switching models is _very_ expensive in compute (you have to rerun everything from the beginning), and highly variable in cost. Cursor tried doing this for awhile, but inconsistent performance&#x2F;usage means most users turned it off and pick models specifically.
    4. condour75 · · focus · HN ↗
      I guess the question is, does the Dunning Krueger effect apply to models? The dumb ones might think they&#x27;re up to the task.
    5. SkyBelow · · focus · HN ↗
      These models have a knowledge cutoff that don&#x27;t just prevent them from knowing about themselves (especially since most data about the model doesn&#x27;t even exist until after the model is created), but they also don&#x27;t know about other recent models. Sure, they can search and use other sources, even make some guesses based on the models they do know, but their default stance is more akin to &quot;User asked about model X, model X doesn&#x27;t exist, maybe it was an hallucination or mistake, let me do a web search...&quot;, but that assumes they have web search and are willing to spend tokens on it.

      Personally I&#x27;ve taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.

      The speed I&#x27;m having to update that document has not gone unnoticed.

    6. Mkengin · · focus · HN ↗
      Github is trying to do that: <a href="https:&#x2F;&#x2F;github.blog&#x2F;ai-and-ml&#x2F;github-copilot&#x2F;project-hydrafusion-frontier-quality-via-multi-model-orchestration&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.blog&#x2F;ai-and-ml&#x2F;github-copilot&#x2F;project-hydrafu...
  7. phpnode · · focus · HN ↗
    What&#x27;s driving the increase in release cadence here? We seem to get new models every week or so now, is this RSI?
    1. jesse_dot_id · · focus · HN ↗
      No.
      1. anotha_one · · focus · HN ↗

        [dead]

    2. lxgr · · focus · HN ↗
      Wanting to have the newer model than the competitor, presumably.
      1. dandellion · · focus · HN ↗
        The old &quot;the bigger number is better&quot;, GPT announces model 6.1, the obvious thing to do next is to announce Gemini 27, and after that Claudé 3000, then a flute album.
        1. lxgr · · focus · HN ↗
          We swear, We Really Wanted To Make An &quot;ASI&quot; Model But This Is Literally The Way The Weights Dragged Us This Time
    3. sharpshadow · · focus · HN ↗
      Response to DeepSeek’s technical paper and competition.
      1. wg0 · · focus · HN ↗
        What&#x27;s that in summary?
        1. Wheen · · focus · HN ↗
          Not the person you&#x27;re replying to, but judging by the emphasis on the cost of cached input tokens in the OP article, I&#x27;d guess it has to do with DeepSeek v4.1&#x27;s KV cache efficiency. It uses &lt;1000 bytes per token, so they&#x27;re able to get 1M token context in under a GB.

          Edit: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;deepseek-ai&#x2F;DeepSeek-V4.1-Flash&#x2F;blob&#x2F;main&#x2F;DeepSeek_V41_Tech_Report.pdf" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;deepseek-ai&#x2F;DeepSeek-V4.1-Flash&#x2F;blob&#x2F;...

          1. ChromeUltron · · focus · HN ↗
            just goes to show that OpenAI in fact did not innovate on a single thing for the better part of a year (one could argue two) and instead keeps immitating what it sees doing others successfully with the tech, all in a very transparent attempt to get people lubed up for their IPO.
      2. LPisGood · · focus · HN ↗
        Which paper are you referring to?
    4. mckirk · · focus · HN ↗
      No, we&#x27;re pacing ourselves to have the time to evaluate the impact each new model could have, obviously.
    5. system2 · · focus · HN ↗
      Chinese model pressure. Many of my SWE friends switched to Chinese models. I also use QWEN and GLM for many of the api requiring projects and dropped OpenAI and Anthropic. The only reason was the cost.

      EDIT: I love getting downvoted by openai and anthropic employees or their bots.

      1. wg0 · · focus · HN ↗
        I can&#x27;t recommend Chinese models enough. My personal favorite is DeepSeek v4.1 Flash but I have tried Qwen 3.8, Kimi 3 and GLM 5.3 which are equally impressive but DeepSeek is the cheapest and fastest regularly hitting 270 token per second.

        And yeah I have worked with Anthropic and OpenAI models, they&#x27;re good but they cost a fortune while Chinese models are already really good at a fraction of the cost.

        1. copperx · · focus · HN ↗
          [delayed]
        2. andybak · · focus · HN ↗
          DeepSeek v4.1 Flash is fascinating and uneven. It&#x27;s way too chatty in OpenCode to be a collaboration partner. I tried dsh-tui which feels comparable to the codex&#x2F;claude tui&#x27;s and it&#x27;s usable. but it seems to be &quot;brilliant and yet stupid&quot; in a way I can&#x27;t quite put my finger on. I&#x27;ve got too much real work to get done to dig into it so until the big boys price me out of the market I&#x27;m back to my $100&#x2F;month deal.
      2. thraway3837 · · focus · HN ↗
        I keep hearing about these Chinese models, but what exactly are you doing with the models and coding? I have a need to fully write code with full tool calling capabilities. Not just methods or functions. I want to be able to prompt a feature and it makes the JIRA ticket, and fully implements it and makes a PR. I don&#x27;t want to babysit it or even read the code. Once it creates the PR, I want it to monitor it for any comments fro Copilot&#x2F;security review and then fix it as necessary.

        Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?

        1. senderista · · focus · HN ↗
          Sounds like your dream workflow could replace you with a PM?
        2. white_dragon88 · · focus · HN ↗

          [dead]

        3. system2 · · focus · HN ↗
          I am using the API only with them for now. But what you describe is nothing compared to Qwen or Mimo. These models are more capable than Opus in general and at a fraction of Opus&#x27;s cost.
    6. jonatron · · focus · HN ↗
      Probably just the singularity, no big deal
      1. anotha_one · · focus · HN ↗

        [dead]

    7. SwabbyNat74 · · focus · HN ↗
      Its a news cycle more than anything, and its ONLY going to get much, much worse. Daily releases, or multiple daily, 30-45, by EOY. Welcome to RSI!
    8. tjwebbnorfolk · · focus · HN ↗
      Competition
    9. colpabar · · focus · HN ↗
      What I don&#x27;t understand is how much people have to say about every single one. Aren&#x27;t we at the diminishing returns stage yet? Is there really that much to discuss?
      1. infamouscow · · focus · HN ↗
        If you look closely at various benchmarks, you&#x27;ll see that often models will improve in certain areas while regressing in others. It suggests we&#x27;re already at the point of diminishing returns.
    10. Aboutplants · · focus · HN ↗
      I do wonder if people switch back and forth between primary models (GPTvsClaude) that it may be a better idea to simply keep releasing updates as soon as possible in order to keep users from bouncing back and forth.
      1. scrollop · · focus · HN ↗
        Probably one of the factors. Signed up to openai pro a few days ago, deciding between openai and anthropic, then sonnet 5.5 was released and am wondering whether I made a mistake.

        Luckily it&#x27;s not a mistake as now we have access to . . . dots.

        (and sol 6.1, it seems)

      2. sockaddr · · focus · HN ↗
        This is it.

        It&#x27;s because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor&#x27;s version bump is the logical thing to do. It has nothing to do with RSI.

        1. vividfrier · · focus · HN ↗

          [dead]

      3. geeky4qwerty · · focus · HN ↗
        jokes on me, I pay for all the subscriptions.
      4. pythonaut_16 · · focus · HN ↗
        Maybe process maturity too.

        Like think about a software org with good CI&#x2F;CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.

        As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.

      5. killingtime74 · · focus · HN ↗
        Of course they do. The real money makers are not subscription users, but the API users and you can just switch with the model selector.
    11. mattnewton · · focus · HN ↗
      Anthropic’s IPO?
    12. franzcoughka · · focus · HN ↗

      [dead]

    13. motoboi · · focus · HN ↗
      New models are distill from the actual unrelease frontier models. They are just giving us better checkpoints.
    14. mynameisjonny_ · · focus · HN ↗
      The initial response to 6 Sol was bad, and Opus 5.5 was definitely winning the public vibes war. Makes sense to rush something out
    15. agluszak · · focus · HN ↗
      They&#x27;re releasing Sol 6.1 because 1. Astra 6.1 got postponed 2. Sol 6 is shitty 3. They have to release _something_ in response to Opus 5.5
    16. orbital-decay · · focus · HN ↗
      Versions is marketing, snapshots&#x2F;minor variations are easy and the number must go up. Release timing is another OAI&#x27;s marketing tactic.

      &gt;RSI

      Recursive improvement doesn&#x27;t imply increased rate, another word for it is &quot;iterative&quot; but it&#x27;s probably too boring for some.

    17. jchw · · focus · HN ↗
      It is the only way to reduce prices while making it look like a good thing.
    18. esafak · · focus · HN ↗
      Productivity is increasing as the models get smarter; we are ascending the singularity. I&#x27;m serious.
    19. toasty228 · · focus · HN ↗
      Opus 5.5 is better than they anticipated, it&#x27;s faster, smarter, cheaper. I&#x27;m about to change provider for claude and I&#x27;m not the only one
      1. copperx · · focus · HN ↗
        It feels like an updated 4.6. It&#x27;s fantastic.
        1. RGS1811 · · focus · HN ↗
          I ran a battery of tests against a couple of simple prompts to check on thoroughness and verbosity of every available Opus, and 5.5 is a lot closer to 5 than people are letting on. 4.6 remains the best in terms of getting to the point and just doing what you ask. I had switched from 4.7 to 5.5 as my main claude model, but started running into the telltale over-interpretation issues of the 5 series, and have switched back. Something in their RL pipeline has made these models consistently worse IMO.
          1. copperx · · focus · HN ↗
            [delayed]
            1. RGS1811 · · focus · HN ↗
              My main grievance is that any gap in specificity in may statement of a task would lead Opus 5 to invent an interpretation to fill the gap, frequently creating lots of extra work for itself in the process, and often deviating from my intent. This would happen even for very simple things. I once asked Opus 5 to fix a failing unit test in CI (something pretty simple), and it went off on a 45 minute expedition (all in one turn), read a boatload of unnecessary files, massively overcomplicated the assignment, etc. It fixed the test but previous models would have handled this much more straightforwardly.

              A common form of this failure is the model picking up on random wordings from earlier in the session (e.g. some comment it made to me in the middle of a response, that I never explicitly endorsed) and then treating these as hard commitments. Or over-interpreting a specific word choice or clumsy phrasing as if it were a &quot;load-bearing&quot; constraint on the task.

              None of this clumsiness would be so problematic if the model didn&#x27;t have such a strong drive toward autonomy. It&#x27;s much like with people: there&#x27;s no shame in not understanding what you&#x27;re being asked to do, provided you ask clarifying questions. There&#x27;s no shame in ignorance if it&#x27;s wedded to curiosity. Benchmaxing has RLVRed curiosity and clarification straight out of these models. It sucks.

      2. copperx · · focus · HN ↗
        [delayed]
    20. denysvitali · · focus · HN ↗
      They&#x27;re pacing the frontier
      1. blmarket · · focus · HN ↗
        and seems like they&#x27;re claiming Sol&#x2F;Opus are not frontier (and only Astra&#x2F;Fable are)
    21. az226 · · focus · HN ↗
      Mature training pipelines, plus ever expanding RL datasets of increased quality, and mega GPU clusters to finish training in a few weeks. Automated safety and reliability testing.
    22. MisterMunchkin · · focus · HN ↗
      Both labs are spying on each other and they get jelly when the other is releasing a new model, so they have to ship something at the same time so they don’t look bad.
      1. dannyw · · focus · HN ↗
        Both labs are spying? Employees hang out at the same bars, have overlapping social circles, etc. Alcohol does what alcohol does.
    23. ChromeUltron · · focus · HN ↗
      no patrick, ~~mayonnaise~~ a point release of the slopbot is NOT ~~an instrument~~ RSI
  8. cmrdporcupine · · focus · HN ↗
    Aka &quot;we made an oopsie last week and released what should have been GPT 6 Terra with the name GPT 6 Sol&quot;
  9. Lapalux · · focus · HN ↗
    It&#x27;s relentless, isn&#x27;t it?
  10. fraywing · · focus · HN ↗
    &gt; GPT‑6.1 Sol matches GPT‑6 Astra at roughly one-fifth of the cost

    Astra is a pretty impressive model. Excited to try this.

    1. gobdovan · · focus · HN ↗
      They have also cut allowances for subscriptions in half. So even in the best case scenario it&#x27;s about 2.5 times cheaper for Codex users. They just seem to have matched Claude Sonnet 5.5 *API pricing*, but from what I see online, it seems Claude Code now has a much more generous subscription allowance.
      1. Tadpole9181 · · focus · HN ↗
        Only for the $100 subscription, correct?
        1. gobdovan · · focus · HN ↗
          Only for the $200 one. The $100 one was already pretty poor value for allowance&#x2F;$. Without the old $200 sub, I wouldn&#x27;t have used Codex.
  11. anotha_one · · focus · HN ↗

    [dead]

  12. Nevin1901 · · focus · HN ↗
    I love free market competition. We&#x27;re getting insane advancements every day. I remember when llms used to cost an arm and a leg for decent intelligence
    1. [deleted] · · focus · HN ↗

      [deleted]

    2. jeffybefffy519 · · focus · HN ↗
      Spotted the person who hasn&#x27;t used these yet....
      1. s3p · · focus · HN ↗
        [delayed]
    3. phist_mcgee · · focus · HN ↗
      That free market is going to explode when the AI bubble bursts.
      1. [deleted] · · focus · HN ↗

        [deleted]

      2. bopou · · focus · HN ↗
        Safe to assume you are actually shorting the market and not just mindlessly babbling about the impending explosion?
  13. glimshe · · focus · HN ↗
    This is great. But maybe part of the motivation is that 6-Sol wasn&#x27;t as good as initially advertised so they needed to tweak it. I felt a clear degradation in quality in some simple refactoring tasks vs 5.6-Sol.
  14. SirMaster · · focus · HN ↗
    Guys, are we slowing down yet?
    1. scottyah · · focus · HN ↗
      Yes, obviously. They&#x27;re both working to make it cheaper, faster, and better at different industries (3d animations, etc). The only direction they are slowing is raw intelligence.
  15. IshKebab · · focus · HN ↗
    I wish they&#x27;d list the environmental cost. My employer has an unlimited AI budget so I don&#x27;t care about using Astra if it&#x27;s just more profit for OpenAI. I care more if it actually uses 5x more energy.
    1. rs_rs_rs_rs_rs · · focus · HN ↗
      I don&#x27;t understand the point of this, why just now when it comes to llms. Why wasn&#x27;t anyone enraged with the environmental costs of kids playing video games. I would not be surprised the environmental cost of that is an order of magnitude bigger than what llms have.
      1. paulryanrogers · · focus · HN ↗
        Considering Nvidia&#x27;s hard shift to crypto and now AI, I doubt videogames are even in the same ballpark.

        How many DCs are devoted solely to gaming?

        1. lp92 · · focus · HN ↗
          Considering many games make use of cloud computing for online play and similar functions, they probably make up a pretty goot bit of global cloud compute capacity. Likely quite a lot less than the big AI players, but not an insignificant amount.
          1. HelloMcFly · · focus · HN ↗
            The energy costs of the cloud computing required for gaming are substantially less in power - not to mention overall demand - than LLMs. Come on, we&#x27;re not in the same energy ballpark here.
            1. rs_rs_rs_rs_rs · · focus · HN ↗
              &gt;The energy costs of the cloud computing required for gaming are substantially less in power

              Yes but it adds up when you consider that just on Steam alone there are 200 million monthly active users.

              1. [deleted] · · focus · HN ↗

                [deleted]

              2. paulryanrogers · · focus · HN ↗
                Steam games can run on less than 4GB of RAM, and per person are likely played only an hour or so a day. AI hyper scalars use far more to serve a single customer.
            2. empthought · · focus · HN ↗
              You&#x27;re not including the physical supply chain energy consumption of distributing video game equipment in this analysis. Nobody ships LLMs to big box stores and tries to sell them to consumers.
              1. paulryanrogers · · focus · HN ↗
                Historically all those boxes and discs were significant, at least if you sum them all for all time.

                Today physical discs are rapidly becoming a niche that newer consoles just won&#x27;t have. PC games are almost never sold physically anymore.

        2. rs_rs_rs_rs_rs · · focus · HN ↗
          &gt;How many DCs are devoted solely to gaming?

          An entire planet. Just Steam alone has one or two hundres million monthly active users.

          1. paulryanrogers · · focus · HN ↗
            The whole planet&#x27;s DCs are not solely serving games. Apparently games take about 300 TWh whilst AI is around 500 TWh.

            All of gaming is in the 300B USD range while just the CapEx of hyper scalars is already over 200B.

      2. lbrito · · focus · HN ↗
        If you recall history past the last 5 minutes, you will remember that people have indeed been enraged with the environmental costs of things for a long time. Its just that AI seems to have induced a mass amnesia, and people tend to forget about what happened pre 2024.
        1. rs_rs_rs_rs_rs · · focus · HN ↗
          &gt; you will remember that people have indeed been enraged with the environmental costs of things for a long time

          Yeah? Show me the big movements against computer gaming.

          1. lbrito · · focus · HN ↗
            There are movements against consumerism and the environmental impacts of industry in general. Greenpeace is over half a century old.

            The differences with AI are: 1) we are starting off (mid 2020s) from a baseline point of already being in a hopelessly shitty situation, past the 1.5C warming target; and 2) Electronics, chips, data centers etc were already a thing for a long time, but industry took _decades_ to ramp up production to pre-AI levels, and these things are used everywhere for a huge number of things. Now we&#x27;re consuming electronics&#x2F;data centers&#x2F;water&#x2F;power at an unheard-of rate, and for a single purpose (AI) with questionable benefits, besides the private interests of a handful of people.

      3. JDups · · focus · HN ↗
        Because people find video games fun, though I suppose there&#x27;s some vocal people that think of them as bad for society. In contrast the AI companies are promising a torment nexus future.

        I&#x27;d be curious as to how much of internet infrastructure is dedicated to gaming though.

      4. IshKebab · · focus · HN ↗
        I don&#x27;t think video games consume nearly as much power. A PS5&#x27;s power consumption is apparently around 200W. That&#x27;s not enough to run even one GPU, let alone the armada it presumably takes to run Astra.

        Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.

        1. rs_rs_rs_rs_rs · · focus · HN ↗
          &gt; non-AI things

          But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?

          &gt; I don&#x27;t think video games consume nearly as much power. A PS5&#x27;s power consumption is apparently around 200W. That&#x27;s not enough to run even one GPU, let alone the armada it presumably takes to run Astra.

          Just Steam has 200 million monthly active users. Add Steam, PS, Xbox, and whole other devices having gpus and I&#x27;m pretty sure you at least 10x the energy consumption of all ai companies.

          1. IshKebab · · focus · HN ↗
            &gt; But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?

            I dunno what you&#x27;re not getting but a GPU to run games is like 200-500W. A GPU cluster to run Astra is probably more like 10kW.

            Also gamers tend not to spin up dozens of other machines to also game for them.

    2. otterley · · focus · HN ↗
      Given that the number one cost of inference is memory and compute, and the incremental cost of each is energy, cost per inference is roughly proportional to energy consumption.
    3. lp92 · · focus · HN ↗
      Did you care about this when it came to your other computing needs? What PC&#x2F;laptop are ypu running and how efficient is that?
      1. codehorses · · focus · HN ↗
        Likely orders of magnitude different, this is a weak whataboutism.
      2. IshKebab · · focus · HN ↗
        Yes I do. I&#x27;ve got a spare desktop that isn&#x27;t too efficient (probably ~100W idle but annoyingly I&#x27;ve lost my power meter) so I don&#x27;t leave it on even though I would like to use it as a server.

        Laptops use very minimal power - you don&#x27;t need to worry about them. If they didn&#x27;t their battery life would suck.

    4. Bolwin · · focus · HN ↗
      Look at Neuralwatt. They report energy usage with every call as well as aggregate statistics.
    5. sergiotapia · · focus · HN ↗
      I want to energymaxx. Every home should have a nuclear generator for free limitless clean energy. Do not energysimp, we want prosperity for all we must energymaxx and invest heavily in solar&#x2F;battery&#x2F;nuclear.
  16. barrenko · · focus · HN ↗
    A night of no sleep this one will be.
  17. Starlevel004 · · focus · HN ↗
    Okay, now price cut 6 Sol (and rename it to Terra again).
  18. dcchambers · · focus · HN ↗
    GPT-6 Sol released a week ago. Shortest model life ever?
    1. cmrdporcupine · · focus · HN ↗
      Taking GPT-6 &quot;Sol&quot; outside behind the shed and giving it a merciful end is about the best outcome possible.

      Huge misstep releasing it.

      1. algoth1 · · focus · HN ↗
        Bro, I&#x27;m a visual thinker
        1. cmrdporcupine · · focus · HN ↗
          Sorry, I&#x27;m GenX. Growing up they showed us &quot;Old Yeller&quot; in the school gym every year like that was some kind of treat.
      2. tandr · · focus · HN ↗
        [delayed]
  19. iamdelirium · · focus · HN ↗
    I wonder if releasing this soon sort of validates the rumor that Sol 6 was just the Terra model they bumped up and slashed the price.

    Then Opus 5.5 caught them off guard and now they&#x27;re actually releasing the correct sized model.

    1. hyperpape · · focus · HN ↗
      If that were true, they’d have axed their margins.
    2. squidbeak · · focus · HN ↗
      Whether it was or wasn&#x27;t, Terra&#x27;s absence shows Sol has replaced it as the new middle model.
  20. [deleted] · · focus · HN ↗

    [deleted]

  21. A_D_E_P_T · · focus · HN ↗
    Looking at the token prices, if this is half as good as 6-Astra for 3D model creation in Blender, it&#x27;s going to be an absolute game changer.

    Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...

    1. ekun · · focus · HN ↗
      How is it with animations?

      I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.

      1. godwinson__4-8 · · focus · HN ↗
        You need to use the Blender MCP. There is an official plugin for this now, so the third party one can be avoided.

        I&#x27;ve only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.

        There are still rough edges of course. But try the MCP out and judge for yourself.

      2. A_D_E_P_T · · focus · HN ↗
        I&#x27;ve only tried animating models in Astra-6, and I was quite impressed! It&#x27;s rarely able to one-shot things perfectly, but it usually gets pretty close.
    2. lukan · · focus · HN ↗
      Have you tried fable? (I did small experiements and was satisfied, but maybe there are reasons to switch?)
      1. A_D_E_P_T · · focus · HN ↗
        No, because I always hit my Fable quota (Max 20x) in 12 hours on simpler tasks, and I&#x27;d hate to need to buy tokens at API pricing.
    3. CuriouslyC · · focus · HN ↗
      From the results of a lot of YouTubers in the space, I think Opus 5.5 is pretty competitive with Astra in 3D. It&#x27;s slightly worse at spatial detail but better at aesthetics and little touches.
      1. A_D_E_P_T · · focus · HN ↗
        Interesting! Can you share an example?
        1. CuriouslyC · · focus · HN ↗
          <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;@stefan_3d_ai" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;@stefan_3d_ai

          A number of others have done game&#x2F;3d video benchmarks but this guy is probably the most prolific.

          1. kroaton · · focus · HN ↗
            That dude is a grifter. His German friend is even worse.
            1. 2yrrr · · focus · HN ↗
              The poster above carries this tone in his posts where he is the divine one. I bet on a lot of stuff he ain’t got a clue what he’s talking about but hopes people like you don’t catch him out.
    4. therealdrag0 · · focus · HN ↗
      After all the hype, I’ve been kinda disappointed tbh. Modeling specific models are so much better (eg. Tripo3d). Astra still models some janky crap for me.
  22. minimaxir · · focus · HN ↗
    &gt; Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing

    This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

    1. TuxSH · · focus · HN ↗
      Exactly half as expensive as Opus 5.5 in every API pricing metric
      1. bigwheels · · focus · HN ↗
        And half as good. I didn&#x27;t have great experiences with Anthropic models in the past, but Opus 5.5 seems to have turned a major corner. It is churning through tasks significantly more quickly and efficiently.

        Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.

        Edit: Defining &quot;difficult&quot; as a complex coding or systems task (or even series of them in a single prompt).

        1. Infinity315 · · focus · HN ↗
          I&#x27;m not an OpenAI simp, but how anyone can have any opinion on the performance of these models in less than a day - let alone a few hours - is beyond me.
          1. toasty228 · · focus · HN ↗
            Try it, it&#x27;s that good compared to openai current offering.

            I get better results and usage our of my $20 claude sub than my $100 openai sub... it&#x27;s that ridiculous

            1. copperx · · focus · HN ↗
              [delayed]
          2. AndrewKemendo · · focus · HN ↗
            Only takes 5-10 minutes to test your favorite one shot comparison prompt.
            1. squidbeak · · focus · HN ↗
              If 5-10 minutes is enough, you need a more ambitious one-shot goal.
            2. edgyquant · · focus · HN ↗
              Can you give an example? For me I find that one shot prompts are pretty good it’s only when working with large codebases and complex, multi prompt workflows, that I find the real limitations of models
              1. AndrewKemendo · · focus · HN ↗
                Yeah the whole Pelican riding the bike is the best obvious one
          3. [deleted] · · focus · HN ↗

            [deleted]

          4. colinhb · · focus · HN ↗
            Yeah totally agree, people keep jumping in w&#x2F; strong views hours after release, eg: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49045430">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49045430
          5. rspeele · · focus · HN ↗
            While I have no experience comparing this brand-new model, OpenAI themselves call it &quot;near-Astra&quot; intelligence. I set Astra and Opus 5.5 independently working on the same large research&#x2F;coding task in an experimental project (doing NURBS surface modeling stuff). They had the same starting repo state, same task packet, same test suite to try to meet. I have the $100 plan in both.

            Astra used 215% of a week&#x27;s budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week&#x27;s budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.

            The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.

            The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn&#x27;t pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus&#x27;s version of the code, but otherwise its approach was behind.

            Now, this is just one comparison in one domain, and arguably Astra&#x27;s strict adherence to the tests as-given is a good thing. But Opus wasn&#x27;t merely loosening the rules &#x2F; moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.

            Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn&#x27;t produce a working implementation (to be fair, Opus&#x27;s was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.

            1. agar · · focus · HN ↗
              This was a very interesting, informative, and well-written comment (and experiment). Thank you.
            2. this_user · · focus · HN ↗
              Astra doesn&#x27;t just burn token at an insane rate, it is also strangely high maintenance when using it. Occasionally, you have to keep prodding it to keep working. Then at other times, it will disappear down some rabbit hole, trying to resolve increasingly hypothetical issues. It feels like you constantly have to keep it on track, while Opus is just churning through tasks.
              1. rrvsh · · focus · HN ↗
                Yes, I really don&#x27;t like Astra - 5.6 models seemed to perform at literally the same level with less opaque prose; I guess Astra is great if you&#x27;re working on insanely hard mathematical problems (or are fooled by its masked sycophancy) but for coding 5.6 seems to have better taste. I hope that they course correct or at least offer models that do better for coding, or even better that this oligopoly ends
            3. chaostheory · · focus · HN ↗
              [delayed]
              1. rspeele · · focus · HN ↗
                I strongly agree!

                My biggest conclusion from this test was: the most efficient use of my weekly Astra budget is as a reviewer&#x2F;consultant for work done by Opus. I don&#x27;t have Astra write much code right now, but I do have it reading a lot of what Opus writes. Of course with the way the AI landscape shifts the balance could be the exact opposite 2 weeks from now.

                Seeing how each model preferred its own flavor of code shows that, even from a &quot;blind&quot; fresh context, a same-model reviewer will still often look at the work of another incarnation of itself and go &quot;yep that&#x27;s how I woulda done it&quot; and not be as likely to realize that there was an alternative path or implicit assumption&#x2F;mistake in the work.

          6. phoghed · · focus · HN ↗
            I think it’s one of the reasons why you often see people decrying the lessening capabilities of the models a few weeks later, despite there being 0 proof of any changes, and evidence of the models staying the same from sites that track it.

            They form these super strong opinions after a few prompts, then face reality over time.

            People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.

          7. beering · · focus · HN ↗
            They’re comparing against the previous model, not the newly released one (6.1). Why do that on a thread about the new model, I don’t know.
          8. ex1fm3ta · · focus · HN ↗
            benchmarks.
        2. mmis1000 · · focus · HN ↗
          For my personal experience, antropic model have better use experience except for 4.7 and 4.8 though. 4.7 and 4.8 feels like expensive downgrade of 4.6 to me (I didn&#x27;t know why these two should even exist)
          1. krzyk · · focus · HN ↗
            For me Anthropic models from 4.7 to 5 including where bad and ate tokens like crazy. Task delivery was worse than GPT 5.6 and token usage was 2-3x higher.

            Looks like 5.5 is the new 4.6

        3. jauntywundrkind · · focus · HN ↗
          A pity I have to use claude code to try this, that I can&#x27;t use the tools I know and love and have built around (opencode).
        4. dotancohen · · focus · HN ↗

            &gt; Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.
          
          That&#x27;s far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I&#x27;ve been finding them much more natural. What is &quot;something difficult&quot; in your workflow?
          1. peterbell_nyc · · focus · HN ↗
            You HAVE to have a set of personal evals for each class of task you want to use models against at scale so you can test plausible candidates and compare output on your work against your evals.

            There is way too much subtlety in what does and doesn&#x27;t work for a given problem, context&#x2F;prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.

          2. Starlevel004 · · focus · HN ↗
            &gt; OpenAI models used to be the prototype for robotic text, but lately I&#x27;ve been finding them much more natural.

            This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.

          3. notatoad · · focus · HN ↗
            My side by side evaluation this week was to build a tool for mounting my app’s UI components in a headless chrome and feeding mock data into them, for the purpose of taking screenshots for help docs. Not super complicated, but a real task I needed done.

            I have the task to codex first, it took a couple back and forth prompts to define the project and then it worked for a bit and to took a couple more prompts before I decided it was good enough - not perfect, but close. It re-implemented some wrapper components in a simplified way that lost some of the UI, but it would work.

            Opus 5.5 took the same prompt with no back and forth, it just went off a built a tool that takes pixel-perfect screenshots of exactly what my app looks like.

        5. TuxSH · · focus · HN ↗
          &gt; Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.

          Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it&#x27;s not as bad as GPT-5.6 Terra I suppose.

        6. sobiolite · · focus · HN ↗
          Are you comparing Opus 5.5 with GPT-6 Sol or GPT-6.1 Sol? Because they are different models.
        7. beering · · focus · HN ↗
          This news and thread is about 6.1 Sol, not 6 Sol. You haven’t even had time to do a fair comparison yet.
        8. rrvsh · · focus · HN ↗
          I found that 6 Sol is dogshit; have you tried o5.5 vs. 5.6 sol? curious to hear if your experience is still the same in that regard
        9. [deleted] · · focus · HN ↗

          [deleted]

      2. dom96 · · focus · HN ↗
        Based on my benchmark[1] it is the same price as Opus 5.5 and just as capable.

        1 - <a href="https:&#x2F;&#x2F;bench.killswitch-lang.org" rel="nofollow">https:&#x2F;&#x2F;bench.killswitch-lang.org

        1. zeroonetwothree · · focus · HN ↗
          Opus 5 scoring higher than 5.5 makes me question of the value of this benchmark to real world usage
          1. dom96 · · focus · HN ↗
            Well, it is genuine.

            Opus 5.5 fails the &quot;understanding&quot; tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1&#x2F;0 instead of the max numbers. Opus 5 gets it right.

            Here are the outputs from both: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;dom96&#x2F;b5bce82b6e6c1ebd5271ed70ad941b49" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;dom96&#x2F;b5bce82b6e6c1ebd5271ed70ad941b....

            Looking at that Opus 5.5 fails to deduce that the &quot;hack statement&quot; is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5&#x27;s greater intelligence for what it&#x27;s worth.

    2. joshstrange · · focus · HN ↗
      &gt; 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

      Cache doesn&#x27;t help you much when you are compacting every 5 minutes...

      I was shocked at how quickly I ran out my $100&#x2F;mo subscription with a single agent (sol medium).

      1. codewithcheese · · focus · HN ↗
        you can config codex to compact at a higher context limit
      2. redox99 · · focus · HN ↗
        If you run out of sol medium with $100 you&#x27;re doing something wrong. Astra destroys your usage, I get 1 day of usage with Astra, but 6 sol is almost unlimited and I only use xhigh.
        1. shimman · · focus · HN ↗
          &quot;You&#x27;re holding it wrong.&quot; Is hardly a retort from a real paying customer having problems with their paid services.

          This is why these companies are struggling to make money, they&#x27;re chastising their customers just like they&#x27;ve been chastising the human race.

          1. trio8453 · · focus · HN ↗
            &gt; &quot;You&#x27;re holding it wrong.&quot; Is hardly a retort from a real paying customer having problems with their paid services.

            It&#x27;s very appropriate in the cases when you&#x27;re holding it wrong. The fact that you&#x27;re paying doesn&#x27;t mean that you can&#x27;t make mistakes or waste resources.

            1. shimman · · focus · HN ↗
              I don&#x27;t find it appropriate at all, especially regarding technology that workers deeply hate and are skeptical of.

              If this is how you want to get people on your side, I can understand why the entire country&#x2F;human race are against these companies.

              1. trio8453 · · focus · HN ↗
                Sides? Hate? This is all very emotional. Try to put the facts down plainly and see how ridiculous it is --

                It&#x27;s a product and if you&#x27;re using it incorrectly, we can either

                1. say so

                2. pretend that you don&#x27;t so to get&#x2F;keep you on &quot;our side&quot;? or not say is because you&#x27;re skeptical or hate it? (how does that last bit even follow logically?)

                How is 2 better in any way for anyone involved?

            2. crossroadsguy · · focus · HN ↗
              [delayed]
          2. Anonasty · · focus · HN ↗
            Literally the prompting and task definition is main variable how LLM&#x27;s performs. There are literally millions of examples of vibe coders and new AI adopters who run out of tokens since they don&#x27;t know how the LLM&#x27;s work.
        2. jorblumesea · · focus · HN ↗
          yeah I use sol constantly and have done maybe $15 of spend in the past week. it&#x27;s solid and cheaper. this is at least 4-5 investigations, prs, whatever per day.
        3. Aeolun · · focus · HN ↗
          It’s only nearly unlimited if you haven’t just used a banked reset. After a banked reset your weekly usage gets cut by about 80% (not the week you need to wait to get your normal limits back though). ChatGPT has given me a really good reason to cancel.
          1. threecheese · · focus · HN ↗
            Can you elaborate? I&#x27;ve been getting great usage out of my $200&#x2F;mo plan, and thought I&#x27;d try a reset (first time) which was expiring just for giggles. Am I going to get only 20% of it effectively?

            I overused Astra in order to drain my weekly, figuring I&#x27;d have the reset. (not wastefully, I did get more work done)

            1. Aeolun · · focus · HN ↗
              I can’t say what will happen to you, but yes, that has been my experience. It is better to wait for your normal full limit to return, because if you use a banked reset you get only 1&#x2F;5th of the tokens but you still have to wait the full week afterwards for it to reset. 20% would be fine if it didn’t also reset the date your normal reset fires.
              1. seunosewa · · focus · HN ↗
                Banked resets do expire if you don&#x27;t use them, so use them anyway.
              2. nkmnz · · focus · HN ↗
                Did you “earn” that reset on a lower tier?
              3. edot · · focus · HN ↗
                Proof? Like, do you have logs or something? Not calling you a liar but this seems not correct based on my usage.
        4. mattkenefick · · focus · HN ↗
          How do you get 1 day of usage with Astra?

          I create a lot, but I can make a full month with Astra on the current Pro plan. What are you doing to spend that much?

          1. redox99 · · focus · HN ↗
            Currently spending a lot of tokens programming the AI for my videogame.

            1 day is kind of generous, it probably lasts like 12 hours of running non stop.

      3. apitman · · focus · HN ↗
        You have a lot of control over compaction, both directly by changing compaction settings, and indirectly by how you structure your codebase&#x2F;docs so agents use less tokens.
      4. antonvs · · focus · HN ↗
        Try Gemini. It’s so cheap I often use my personal AI Pro account for corporate work, and most of the time it doesn’t matter.
        1. ChickeNES · · focus · HN ↗
          Gemini is dumb as hell though, it&#x27;s not like for like
          1. Foobar8568 · · focus · HN ↗
            cheerleader hallucinating agent. That&#x27;s Gemini.
          2. Marha01 · · focus · HN ↗
            Gemini 3.8 Flash is actually pretty good.
          3. antonvs · · focus · HN ↗
            I doubt you&#x27;ve tried it recently, or perhaps you confused the search engine version of Gemini for the frontier models.

            I&#x27;ve been using Gemini on development of a DNN training pipeline, and there&#x27;s no way you can describe it as &quot;dumb as hell&quot;. That description just makes it clear that you&#x27;re not talking about the technical capabilities of the models, but about some sort of fanboy comparison from a parallel hype universe.

      5. _davide_ · · focus · HN ↗
        As a reference i burn 1% percent for every 40 minutes of sol on average
      6. onlyrealcuzzo · · focus · HN ↗
        If you&#x27;re compacting every 5 minutes, you have a workflow problem - period.

        No LLM will be cost effective if it&#x27;s compacting this often. You have to find a way around it.

        1. ngruhn · · focus · HN ↗
          Context window is only 275k or something. And honestly compaction is not that bad in Codex. I often don&#x27;t even notice I went through 5 compactions in a session.
          1. sally_glance · · focus · HN ↗
            Same for me, I started wondering if maybe workflows using compaction instead of clear + markdown memory would be more efficient. Writing a plan or tasks to a file often has the next session repeat part of the exploration, compaction seems to keep most relevant context.
          2. SyneRyder · · focus · HN ↗
            Sounds like that&#x27;s the problem then, 275k is a tiny context window. I regularly have sessions that go to 450k or even up to 700k for an unattended overnight Claude Opus session.

            Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:

            <a href="https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2089082893804896524" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2089082893804896524

            There&#x27;s at least a forum thread about it here:

            <a href="https:&#x2F;&#x2F;community.openai.com&#x2F;t&#x2F;why-does-codex-report-a-258-400-token-context-window-for-gpt-5-6-sol&#x2F;1394346&#x2F;5" rel="nofollow">https:&#x2F;&#x2F;community.openai.com&#x2F;t&#x2F;why-does-codex-report-a-258-4...

            1. gf000 · · focus · HN ↗
              But that 250k context worth way more than 1M in terms of how well it&#x27;s utilized, so actually I do like codex trying to keep you at that sweet spot.
            2. rrvsh · · focus · HN ↗
              Its really not tiny; you can&#x27;t compare Claude to GPT, they have honestly diverged enough that as the other reply said, 256k GPT is about equal to 1M Claude. The compaction is slightly annoying, and you can turn it up to 1M as you said if you truly need everything in context, but otherwise it&#x27;s perfectly serviceable
              1. kaoD · · focus · HN ↗
                [delayed]
            3. bjord · · focus · HN ↗
              &gt; unattended overnight claude opus session

              yes, exactly

              1. SyneRyder · · focus · HN ↗
                Not sure I understand if this was meant as a slight against Claude? Or agreement?

                These are often my best sessions - they&#x27;re unattended overnight, because by then we have the specification figured out, and I can just leave Claude to build out the rest, making good choices if it does find gaps in the spec. I regularly go to sleep &amp; wake up to an entirely new application completed. Claude never uses compacting in my sessions.

                I haven&#x27;t used GPT as much as I should have, so I&#x27;m prepared to be incorrect &amp; out of date. It just intuitively feels like I wouldn&#x27;t get the same from a 275K context window - maybe it uses lots of subagents? Even Deepseek &amp; GLM have 1 Million context windows now, so it &quot;feels&quot; strange for people to actually prefer the 275K window. But that&#x27;s just my intuition.

                1. bjord · · focus · HN ↗
                  neither, actually, just that unattended &quot;oneshot&quot; sessions are incredibly token inefficient

                  if you talk about them (in which you lean on an LLM as a sort-of independent employee) and conservative, chunk-based usage (in which you use the LLM as more of an extension of yourself), you&#x27;re comparing apples to oranges

                  a predefined spec obviously reduces that gap but how much is highly dependent on the level of detail

          3. jeremyjh · · focus · HN ↗
            I don’t usually have a problem doing a complete task in that context size. OMP does make a lot of use of rewind which may be helping - basically forks itself and sends back a summary after a long tangent. Coding takes use a Luna max agent.

            I’ve also found compaction not to be a problem when it does happen.

            1. threecheese · · focus · HN ↗
              How do you trigger this? I&#x27;ve been messing with OMP lately for funsies.
              1. jeremyjh · · focus · HN ↗
                Its in settings under Tools-&gt;Checkpoint&#x2F;Rewind. I don&#x27;t know why its not enabled by default.
          4. onlyrealcuzzo · · focus · HN ↗
            If it&#x27;s compacting every 5 mins, you&#x27;re going to notice it in your cache miss ratio and your costs...
            1. rrvsh · · focus · HN ↗
              It doesn&#x27;t - try it first
          5. Benjamin_Dobell · · focus · HN ↗
            The context window is configurable. I&#x27;ve been using ~600k for months. No, not API pricing, on a Codex sub.

            ~&#x2F;.codex&#x2F;config.toml

              model = &quot;gpt-6.1-sol&quot;
              model_context_window = 700000
              model_auto_compact_token_limit = 630000
      7. AmazingTurtle · · focus · HN ↗
        you can actually leverage 400k and 1M contexts in codex with very little code changes to the harness. note that excess context past the.. 250k or 400k mark (i don&#x27;t remember) is charged at 2x the price.
      8. manmal · · focus · HN ↗
        Your tool calls (MCPs?) are very likely too wasteful. Apply some filtering logic.
      9. Gareth321 · · focus · HN ↗
        &gt; Cache doesn&#x27;t help you much when you are compacting every 5 minutes...

        It&#x27;s crazy on Codex. I sometimes get just 2-3 turns before it compacts. It has forced me to use persistent project documentation for everything. Maybe that&#x27;s not a bad thing but unless it reads all the documentation after every compaction (and uses half its cache), it goes off the rails. By comparison, Opus 5.5 is a breath of fresh air. It takes FAR longer to hit the cache limit and that means it keeps useful information in working memory far longer. I think this alone has resulted in a massive productivity and efficiency increase for me.

        1. RugnirViking · · focus · HN ↗
          iirc you can still turn the compaction limit up in codex, though they don&#x27;t make it easy. It costs way more when you use &quot;large context&quot; though, more than the ~256k that codex allows by default. You can also use the large context via the api directly
      10. exfalso · · focus · HN ↗
        what. I use Astra xhigh, sometimes max, never ran out of tokens on the 100$ thing. I&#x27;m using pi though which is by definition harder better faster stronger than claude code&#x2F;codex.
      11. jmalicki · · focus · HN ↗
        Use more subagents.

        The longer your chat gets, the slower and more expensive it gets.

        Subagents are expensive but they scale way closer to O(n) than O(n^2).

        Have some agents make bug reports&#x2F;feature requests&#x2F;roadmaps (linear is very AI friendly), others coordinate, others work on grinding out an individual ticket.

        If there is a good ticket-level description, it&#x27;s a waste of time IMO to have a main agent do it, that should be an agent with fresh context that will do it better faster (the shorter the context, the better models are at using the context they&#x27;re given).

        1. jaktet · · focus · HN ↗
          Subagents will inherit the context window at the point in which they are spawned, but it sounds like you&#x27;re more referring to orchestrating&#x2F;conducting&#x2F;managing multiple agents?
          1. jmalicki · · focus · HN ↗
            Both... even subagents inheriting the context window doesn&#x27;t cost a huge amount if the context window was never that large, but yes orchestrating&#x2F;conducting&#x2F;managing multiple agents is even better though higher thought cost (but the newer claude agents are really good at this in my experience, part of why I am using Claude a lot lately despite the models being more expensive that ChatGPT&#x27;s for the same performance when taken alone).

            Whenever I see my main agent do a compaction, that to me is a clear sign I didn&#x27;t have it delegate bounded tasks enough.

            Still, I see no evidence Codex or Claude Code inherit full context of the main agent in subagents, I&#x27;ve always seen them be prompted, but this is something high priority on my list of unknowns to understand better...

      12. KetoManx64 · · focus · HN ↗
        Do you just keep one conversation going for all projects? That&#x27;s the only way I&#x27;ve seen other people in my company burn through all their tokens.

        Everyone else that uses memory files and a new conversation for each new sub project&#x2F;feature rarely hit their weekly allotments.

    3. verdverm · · focus · HN ↗
      cache is typically 10%, is this OAI setting a new level at half, 5%?
      1. crazylogger · · focus · HN ↗
        The backdrop being deepseek offering 1% (I remember it was ~1% when 4-pro first came out early this year - 4-pro is now removed) &#x2F; 2% (current for 4.1-flash).
    4. sscaryterry · · focus · HN ↗

      [dead]

      1. JimDabell · · focus · HN ↗
        &gt; most people, get hardly a days usage out of a 20x account

        This is not even remotely true.

        1. peterbell_nyc · · focus · HN ↗
          This is the distribution of usage. Spin up a bunch of loops or fire a semi-autonomous factory at a project and it&#x27;s pretty easy to blow through a 20x account in a few hours if you can afford the sandboxes, CI and other infra required.

          If you&#x27;re running 2-3 parallel agent session with a few sub agents and waiting for you to prompt them, you&#x27;ll have a very different experience!

          1. JimDabell · · focus · HN ↗
            &gt; Spin up a bunch of loops or fire a semi-autonomous factory at a project and it&#x27;s pretty easy to blow through a 20x account in a few hours if you can afford the sandboxes, CI and other infra required.

            This is a tiny minority of people, not “most people”.

      2. user43928 · · focus · HN ↗
        It&#x27;s obviously true.

        With the 80% price cut, this is competitive with Opus 5.5 despite the subscription downgrade.

        Additionally, it was said that existing 20x subscriptions retain the higher limits for some time.

        I have seen you make these immature accusations that users here are OpenAI employees multiple times today.

        1. sscaryterry · · focus · HN ↗
          It is not obviously true. Please provide real proof. OpenAI&#x27;s customers are tired of their BS.
      3. minimaxir · · focus · HN ↗
        if an openai employee is reading this plz hire me i am unemployed and i need a job

        (Usage limits are entirely dependent on what you&#x27;re doing with them. If you&#x27;re not running it on 1 million LoC databases you can get a lot of mileage out of a 5x account)

    5. vcryan · · focus · HN ↗
      People&#x27;s volume and approach varies. I&#x27;m a happy customer and I use my entire double max subscription on planning and analysis and have other models doing all my implementation work because I would burn through my subscription in a day or less. It&#x27;s difficult to calculate, but I&#x27;m something like 10-20 billion token per week consumer and I can&#x27;t use a US-based model to do this volume of implementation work.

      Also, a lot of this work is verification to ensure that AI generated code does what is intended and is safe to merge and deploy. That verification work is critical and uses a lot of tokens.

    6. pvab3 · · focus · HN ↗
      gpt 6 Sol was already supposedly better and 50% cheaper than 5.6 Sol right? I didn&#x27;t understand why they were keeping 5.6 Sol
  23. throwitaway222 · · focus · HN ↗
    So yesterday we were consumed with how this was being delayed because of safety, yada yada.

    Guess not?

    1. murbard2 · · focus · HN ↗
      Astra 6.1, this is Sol 6.1
    2. minimaxir · · focus · HN ↗
      That model was implied to be GPT 6.1 Astra, not Sol.
    3. nimonian · · focus · HN ↗
      I see where you are coming from. But 6.1 Sol seems like a new frontier in pricing, not intelligence. I do think the deceleration stuff was mostly bluster, but I don&#x27;t think this release in particular contradicts it too much.
    4. anotha_one · · focus · HN ↗

      [dead]

  24. zf00002 · · focus · HN ↗
    I can blow through my weekly on astra in a few hours; hopefully this really is as good.
    1. skybrian · · focus · HN ↗
      [delayed]
  25. alvis · · focus · HN ↗
    Cache is priced at $0.1&#x2F;M, 50% as sol 6 and sonnet 5.5.
  26. [deleted] · · focus · HN ↗

    [deleted]

  27. slopinthebag · · focus · HN ↗
    $2&#x2F;10 is pretty cheap for a frontier model...
  28. sehw · · focus · HN ↗
    I stopped using LLMs. I shit you not. My life got better.
  29. Aboutplants · · focus · HN ↗
    So when does Anthropic answer? Tomorrow?
    1. Alifatisk · · focus · HN ↗
      You live in a ping-pong.
    2. iosjunkie · · focus · HN ↗
      hopefully the answer doesn&#x27;t include an increase in cost&#x2F;decrease in usage.
  30. mekpro · · focus · HN ↗
    Why they are not even benchmark model against Anthropic or anybody ?
  31. jdw64 · · focus · HN ↗
    GPT 6.0 Sol was so terrible—I wonder if 6.1 Sol will be good?
  32. godwinson__4-8 · · focus · HN ↗
    Let&#x27;s all boycott and move to Claude until they release 6.1 Astra. I don&#x27;t like to be teased.

    When is the alleged &quot;safety&quot; concern satisfied? Does this mean releasing new capability to consumers is going to get a lot slower? Lower price for 6 Astra capability via this 6.1 Sol is exciting, but that is because of Astra capability not merely the low price point.

    When do we get the next jump in capability? When is 6.1 Astra released?

    1. ColonelPhantom · · focus · HN ↗
      Isn&#x27;t Anthropic doing the same, with Opus 5.5 being out while Fable&#x2F;Mythos is still on 5.1?
      1. godwinson__4-8 · · focus · HN ↗
        Is this due to a similar safety concern or just because it&#x27;s not ready yet?

        The coverage around 6.1 Astra seems deliberately playing into the dubious &quot;safety&quot; narrative in a way that feels distinct. But you may be correct in which case, I would take the correction on board and maybe suggest a different alternative.

        Although in theory if OpenAI was boycotted in this way the market pressure would force them to release. Then everyone moves over there. Then Claude faces the same pressure. So even so, I think it could still work even if you have to trade off who you are boycotting from time to time.

        Without more details on the credibility of the safety concern this seems like a totally coherent action for customers to take. We shouldn&#x27;t put up with teasing.

      2. wren6991 · · focus · HN ↗
        It&#x27;s just vibe versioning, right? Fable 5 is a beloved product, it gets a .1 bump to feel close. Opus 5 and Sonnet 5 had a mixed reception, they get a .5 bump to create a sense of distance.
  33. vb-8448 · · focus · HN ↗
    The real announcement is the ultra fast mode ... Astra at 300t&#x2F;s is insane!
    1. tandr · · focus · HN ↗
      [delayed]
      1. az226 · · focus · HN ↗
        6x
      2. TuxSH · · focus · HN ↗
        8x for subscription, 6x for credits&#x2F;enterprise: <a href="https:&#x2F;&#x2F;learn.chatgpt.com&#x2F;docs&#x2F;agent-configuration&#x2F;speed" rel="nofollow">https:&#x2F;&#x2F;learn.chatgpt.com&#x2F;docs&#x2F;agent-configuration&#x2F;speed
    2. objektif · · focus · HN ↗
      Where do you see this?
      1. vb-8448 · · focus · HN ↗
        i&#x27;m watching the keynote
  34. tultra · · focus · HN ↗
    Still not available to me
  35. alexandroskyr · · focus · HN ↗
    Are you guys in dev day? did they start?
  36. formvoltron · · focus · HN ↗
    didn&#x27;t 6 sol just come out a couple weeks ago?
    1. Alifatisk · · focus · HN ↗
      Other comments have already addressed this.
  37. gavin_gee · · focus · HN ↗
    the race to the bottom on models is well underway. huge IPO&#x27;s only really make sense for DC&#x2F;HW lockups, and going vertical.
  38. nicce · · focus · HN ↗
    It is hard to trust these scores. GPP 6 Sol has been so bad for few days.
  39. thefounder · · focus · HN ↗
    They need to fix Astra first. My main issue is with GPT in general is that unless steered it goes into AI sloppiness&#x2F;machinery that is not “needed”. The good part is that this kind of behaviour also makes it good to find subtle bugs or debug issues that Fable&#x2F;Claude just cannot get&#x2F;fix even when you point it.
  40. the_duke · · focus · HN ↗
    The GPT 6 release was ... not great.

    Sol 6 was so bad that I switched over to Opus 5.5 exclusively.

    Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.

    Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.

    I&#x27;m a bit sour on OpenAI right now and skeptical that 6.1 will be much different.

    (Note: this is after preferring and shilling Codex&#x2F;OpenAI models for the last half year)

    1. nxc18 · · focus · HN ↗
      How does this jive with the exponential growth claims? Theoretically sol models are better than the 4 series models I was using at the beginning of the year, but in practice the results don’t seem to be much better. They always nerf the models over the course of the release so it _looks_ like the next version is better but I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling&#x2F;harness improvements.
      1. user43928 · · focus · HN ↗
        They never nerfed any model after release.

        The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.

        I am very skeptical of claims that old models weren&#x27;t much worse. Compare this to February&#x27;s GPT-5.3.

        1. nxc18 · · focus · HN ↗
          I am comparing to GPT-5.3 and 5.2, and I perceive that things have not been noticeably better since then. I also know that I can predict new model releases with high accuracy when my coding agent suddenly becomes regard-level at following instructions and completing simple tasks. This is how I knew 6.0 was about to be released - 5.6 suddenly got unusably bad.

          I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.

          1. holbrad · · focus · HN ↗
            I&#x27;ve heard very little positive press around Sol 6, with a ton of people preferring Sol 5.6 instead.

            I haven&#x27;t used it much yet, but I have much higher hopes for Sol 6.1, as it seems to be based off of a completely different base, it&#x27;s not just a tune.

          2. sebzim4500 · · focus · HN ↗
            I never used sol-6.0 in part because everyone kept talking about how bad it was.

            Astra is clearly far better than anything prior though, so I&#x27;m not sure what you mean really.

        2. Hammershaft · · focus · HN ↗
          <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;

          Am I misinterpreting this, or did OpenAI clearly nerf GPT-6 Sol on the 23rd.

          1. user43928 · · focus · HN ↗
            It released on that date, the data before is a different model.

            The chart shows GPT-5.6 Sol and a surprisingly large drop in performance when the switched it over to GPT-6 Sol.

      2. sigbottle · · focus · HN ↗
        How large of codebases are you working on? The models have gotten good enough to 1 shot stupid &quot;trivial&quot; throwaway integration projects with 0 handholding (was having RL&#x27;d garbage in late 2025), and I&#x27;m actually enjoying designing bounded greenfield personal software from scratch with Astra, in my experience. It&#x27;s quite slow - 2 weeks of credits and constant talking and back and forth with Astra, but it doesn&#x27;t feel annoying to talk to and is like an intelligent colleague maybe 70% of the time? Which is great. Just push back when it&#x27;s dumb.

        I&#x27;m by no means an AI booster, but given 2022 - 2026 progress I&#x27;d say it&#x27;s &quot;exponential&quot; in the sense of, &quot;holy shit, every year I can do more and more genuinely different things&quot;, not &quot;RSI mind reading intelligence can do anything is here&quot;.

        I don&#x27;t think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.

        &gt; I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling&#x2F;harness improvements.

        Even if that were the case, I&#x27;d say that it&#x27;s improved in practice. And just from a philosophy perspective, if you&#x27;re trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).

        1. nxc18 · · focus · HN ↗
          It’s 50&#x2F;50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).

          On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.

          5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.

          1. sigbottle · · focus · HN ↗
            &gt; It’s 50&#x2F;50 on whether it will fuck up implementing an integration test suite when given a list of tests to write and examples of existing tests. It still adds needless abstractions (the reference count codelens in VS Code is good for detecting this sort of thing).

            Yes, still running into this, but surprised about this

            &gt; On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.

            I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just &quot;grasp&quot; the right level of &quot;here is the essence of what we need&quot; versus &quot;these are all the small impl details&quot;. But idk I feel like Astra&#x27;s the first model in quite a while that I don&#x27;t feel genuinely annoyed at handholding a toddler with a PhD.

            But I totally believe you on the 50&#x2F;50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of &quot;make user retry in this case&quot;, it silently built an extremely elaborate recovery state machine w&#x2F;o looking. These pathologies by no means gone, and I&#x27;m still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra&#x27;s gonna do this kind of RL slop failure mode.

            For my use cases personally though, it&#x27;s been better and better. I can&#x27;t use AI at work, so you have much harier edge cases than I do, but still.

          2. moshegramovsky · · focus · HN ↗
            This is 100% absolutely my experience as well. Especially the needless abstractions and endless rounds of corrections. That was literally my entire last week of work.
        2. moshegramovsky · · focus · HN ↗
          I work on a very large code base (millions of LOC) and I&#x27;ve had lackluster results with autonomous work and 1 shotting. AI is definitely fantastic at working on many programming problems but I am not seeing amazing results at refactoring. In fact, I am seeing very poor results, even with Astra, even with extensive planning docs. All the recent models I&#x27;ve used can definitely get that refactor done, but not autonomously. It needs to be small slices. I&#x27;ve yet to see it 1 shot anything really complicated.

          Here&#x27;s a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That&#x27;s a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn&#x27;t. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn&#x27;t follow existing API patterns.

    2. jstummbillig · · focus · HN ↗
      Eh. What? Is this common sentiment?

      I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?

      (Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)

      1. the_duke · · focus · HN ↗
        On r&#x2F;codex the sentiment seems to be quite wide-spread.
        1. phoghed · · focus · HN ↗
          <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;

          Codex itself seems to have a regression. You can see clearly the token use changing significantly coincides with a score drop

      2. nicce · · focus · HN ↗
        When GPT 6 Sol &amp; Luna were released, everything went down. I have been running Sol at max thinking and it is about the same as old Luna with max thinking, give or take. Sometimes feeling even dumber. I can&#x27;t trust it to do anything big alone anymore without babysitting.
      3. copperx · · focus · HN ↗
        [delayed]
        1. Marha01 · · focus · HN ↗
          They should etch it into an ASIC. The first model worthy of that honor.
      4. Eridrus · · focus · HN ↗
        Sol 6 definitely feels kind of dumb and worse than 5.6

        Astra seems better though.

        Showing one potentially saturated benchmark doesn&#x27;t necessarily fill me with a lot of confidence in the coding results.

      5. phoghed · · focus · HN ↗
        In my experience, no. There’s no way to know though. The whole conversation and industry are a combo of benchmaxing, faith, and mysticism.

        Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.

    3. btbuildem · · focus · HN ↗
      That mirrors how disappointing Opus 5 and Fable were, for anything beyond one-shotted tasks or shiny demos. Maybe OAI is just a step behind Anthropic? Opus 5.5 seems like the real deal again, consistent good results on large, complex codebases.
    4. wkcheng · · focus · HN ↗
      I agree, and I haven&#x27;t seen other people mention this! The benchmarks for GPT 6 Sol are great, but realistically it does not seem better than 5.6 Sol. 6-Sol is noticeably worse for code reviews (worse than Deepseek 4.1 flash), has implementation issues (requires more rounds of code reviews and fixes to get to a serviceable state). Opus 5.5 is much much better.

      I&#x27;ve implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.

      If 6.1 Sol has actually matched Opus 5.5, I&#x27;d be very happy. However, benchmarks and real usage don&#x27;t seem to agree in my own tests. So we&#x27;ll have to see.

      1. equinumerous · · focus · HN ↗
        If the benchmarks show better performance, but a consensus of experienced software engineers establishes that the model is worse on coding performance... well, the benchmarks don&#x27;t mean much, do they? It seems like we need much more comprehensive and better benchmarks. And of course, I don&#x27;t think benchmarks yet capture the &quot;human&quot; factor - does a human think a bit of code is logical and maintainable? I often find that these models produce a bit of code, but it is much more convoluted than it needs to be. It makes perfect sense given that these things are code generators, that they generate a lot of code. But quantity of code does not mean code quality, and code quality tends to matter when you read code much more than you write it.
    5. trentnix · · focus · HN ↗
      That&#x27;s not been my experience. My experience with Astra (I use it at home writing Go and C) for coding has been fantastic. Opus 5.5 (I use it for work writing C#) seems faster, but it doesn&#x27;t seem demonstrably better and is still prone to word vomit.
      1. chronogram · · focus · HN ↗
        Same here. Astra has been the best thing I&#x27;ve seen. Astra on Low has been my favourite thing so far. Higher levels just mean more cruft, not useful.
      2. r0l1 · · focus · HN ↗
        Made the opposite experience. Astra was not good in writing go and c++ code. Had multiple OpenAi and Claude subscriptions and all our coworkers agreed. Switched back to Claude and the experience is so much better. Not vibe coding, but assisted coding with immediate feedback.
    6. setnone · · focus · HN ↗
      yeah i can relate, sol 6 is definitely dumber than 5.6, lazier too, i hope it&#x27;s just roll out pains
    7. ozgung · · focus · HN ↗
      Maybe OpenAI was the only one pacing the frontier.
    8. bitexploder · · focus · HN ↗
      I have likewise not been impressed with Astra 6 for most things. It is good, but Opus 5.5 seems just as good or better and I have had Opus 5.5 workers just... hammering since release and cannot spend all of my quota yet.
    9. NorthSouthNorth · · focus · HN ↗
      I shilled so hard to a friend that he actually swapped decided to swap over to Codex. I feel a bit guilty now lol (tbh Astra is a great model, but 5.5 is just brilliant).
    10. sunaookami · · focus · HN ↗
      gpt-6-luna is terrible. It leaks tool calls and markers in the output like crazy, there is definitely something wrong here. gpt-5.6-terra works fine. Also, gpt-6-luna was sneakily added to the 1 mio free tokens group instead of 10 mio. like gpt-5.6-luna: <a href="https:&#x2F;&#x2F;help.openai.com&#x2F;en&#x2F;articles&#x2F;10306912-sharing-feedback-evaluation-and-fine-tuning-data-and-api-inputs-and-outputs-with-openai#:~:text=What%20models%20are%20included%20in%20this%20offer" rel="nofollow">https:&#x2F;&#x2F;help.openai.com&#x2F;en&#x2F;articles&#x2F;10306912-sharing-feedbac...
    11. stldev · · focus · HN ↗
      My experience as well.

      For coding specifically, I&#x27;ve found 5.6-Sol &gt; 6.0 Sol &gt; Astra.

      For modeling and artwork, Astra has been great routinely outperforming Kimi.

      This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.

      I can&#x27;t wait for technology to catch up to a point where we can rid ourselves of this oligopoly.

      1. rrvsh · · focus · HN ↗
        Hard agree

        I had to switch back to 5.6 Sol after trialling 6 Sol for like 3 days - I was getting insanely annoyed at how misaligned it is. Will try 6.1 but not very high hopes

      2. keyle · · focus · HN ↗
        I&#x27;d even go one more, 5.4 was great until 5.5, which was a rug pull.

        I&#x27;ll just leave this here: <a href="https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;" rel="nofollow">https:&#x2F;&#x2F;marginlab.ai&#x2F;trackers&#x2F;codex&#x2F;

      3. dannyw · · focus · HN ↗
        Adaptive thinking was a good idea though. The old method of manually specifying how many thinking tokens you wanted as budget was just silly. The rollout might not have been great, but the change is good.

        And, is it really even an oligopoly anymore? Open weight models are incredibly competitive in every way; whether you want to use US providers, Chinese official providers, self host, etc.

    12. moshegramovsky · · focus · HN ↗
      100% hard agree.

      I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30&#x2F;40 minutes) to do simple things. As a result, Astra didn&#x27;t get much done. It needs the same small implementation slices as GPT 5.5&#x2F;others, but was much slower and didn&#x27;t generate better results. (On a complex infra project&#x2F;across a large codebase.)

      It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn&#x27;t seem to be better than 5.5 at most programming jobs.

      I&#x27;m on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can&#x27;t get coffee fast. Like I can&#x27;t send an email fast.

      OpenAI is making some excellent products for sure but I&#x27;m not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren&#x27;t good at working autonomously on large codebase situations. Just because something compiles doesn&#x27;t make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe&#x2F;side app and then worked through the design there. Um, what? Not that it&#x27;s invalid to do this but I actually have to test in the live codebase or I can&#x27;t possibly say that something is working.

      Just because you can, doesn&#x27;t mean you should.

    13. jrflo · · focus · HN ↗
      I&#x27;m in the same boat, I&#x27;ll give 6.1 a shot but I&#x27;ll probably hop over to Anthropic now that the $200 tier has equivalent weekly usage between the two of them.
      1. skeptic_ai · · focus · HN ↗
        Anthropic 20x plan only refers to 5h interval. Not the weekly quota. Very sketchy
    14. soulofmischief · · focus · HN ↗
      I have had the same exact experience. I feel like I&#x27;m working with 5.3 again. It is alarming how degraded the experience has become over the last month.

      What was a pleasant and productive experience is becoming increasingly frustrating and draining.

    15. beebmam · · focus · HN ↗
      gpt-5.6-sol is significantly better than gpt-6-sol. Not impressed with this new line.
      1. diego_sandoval · · focus · HN ↗
        Agree.

        GPT 6 needs to be babysit, otherwise it starts doing ridiculous things.

      2. 4b11b4 · · focus · HN ↗
        Didn&#x27;t even bother trying it yet
    16. jsw97 · · focus · HN ↗
      After seeing a number of hit or miss releases from both OpenAI and Anthropic my default is to stay put on what I’m using and then free ride on discerning eager adopters by reading their reviews. (Thanks!) Still on sol 5.6 with an occasional advice from Astra. Also I feel like I kind of get used to the models but maybe that’s just my imagination.
    17. pampas · · focus · HN ↗
      That&#x27;s my experience too. GPT-6 Sol tends to rabbit hole and over engineer things.
    18. jeffybefffy519 · · focus · HN ↗
      Its almost like the &quot;frontier&quot; is a load of marketing bullshit and we should ignore it....
    19. koyote · · focus · HN ↗
      I think the fact that Sol 6 appeared higher than Sonnet 4 on benchmarks shows that benchmarks are completely rubbish and useless.

      I&#x27;ve never seen such a large degradation in intelligence in a model until I tried out Sol 6 after having used 5.6 almost exclusively for several weeks.

    20. twotwotwo · · focus · HN ↗
      I am always uncertain about impressions, but mine agree with this. I liked Luna 5.6 on Amazon Bedrock (which got &gt;100 tps) for doing well-specced tasks fast. 6 seems to both be served slower by Bedrock and may spend more turns&#x2F;tokens to get to the same place, so...not as fun.

      And, of course, GPT-6 came out as Anthropic fixed a bunch of stuff with their models -- faster (via fewer tokens, and TPS for Sonnet), easier to work with, better results, cheaper (via pricing and, again, fewer tokens). I don&#x27;t know if the timing and the sudden improvement on Anthropic&#x27;s side sharpened the vibes comparison this round, but Internet opinion went pretty clearly to Anthropic.

      FrontierCode&#x27;s results make it look like the lower two effort settings of Sol-6.1 may be better options than Sonnet or Opus on low, but you might be better off with Opus than with Sol&#x27;s high settings.

      One thing I don&#x27;t think any of this reflects is that many well-specified coding tasks, including the self-testing and doing research and tracing out dependencies and so on, aren&#x27;t really bleeding-edge now: Luna-5.6 and small open models handle them fine. Stuff like &quot;why is this box dropping connections?&quot; or &quot;here&#x27;s a thing I want you to model&#x2F;figure out&quot; can benefit from bigger models. But far from everything does!

    21. laurels-marts · · focus · HN ↗
      100% in agreement. I pay for OAI sub and also use Codex exclusively at work for the past 8 months.

      I tried out fable 5.1 the day it was released and coming from gpt-5.6-sol I was truly mind blown (both in terms of code and prose it was generating - outputs I could finally enjoy reading and looking at).

      Then when opus 5.5 came out, again same thing + far cheaper and faster.

      I went from using OAI exclusively the entire year to a point now where i haven’t touched one of their models in at least a few weeks now.

      I think OAI has lost the plot. OAI models simplify have no taste. And I don’t mean in front-end design way (although that too). They have no taste in how the model writes code, how it writes prose, how it writes in-line comments, how it writes documentation, or how it even picks variable names. There’s just no taste throughout.

      Anthropic models are very thoughtful and have so much taste all around.

      1. stasomatic · · focus · HN ↗
        I cancelled Claude because of its thoughtful prose. I prefer one liner responses from OAI models.
    22. jp_gorman · · focus · HN ↗
      yup - It was darn rude (as the Aussies would say)... Was taking forever to do anything and it made a mess of a lot of resourcing I was working on for som Dev Ops work I was doing on an AI project. Switched back to 5.6 sol and it immediately noted all the mess of worktrees and PRs it had left all over the place.
    23. jp_gorman · · focus · HN ↗
      Fully agree - had to drop to 5.6 sol as 6.1 sol was atrain wreck in my project leaving dead boddies everywhere... 5.6 sol spotted it all the minute it looked. 6.1 sol was noticably major slow down also.
  41. epolanski · · focus · HN ↗
    DeepSeek and GLM made it impossible for &quot;sota&quot; to price any way they want.
  42. modeless · · focus · HN ↗
    GPT 6 Sol is obsolete after only one week! I am glad that they are updating the model more frequently.
    1. pmdr · · focus · HN ↗
      I had it write some code the other day, boy was it awful-looking compared to 5.6. Worked perfectly, but ugly nonetheless.
      1. samuelknight · · focus · HN ↗
        Sol 6 was a flop. Nobody would have cared if it was called Terra 6.
  43. mkaic · · focus · HN ↗
    I got a popup in my Codex just now saying &quot;Try out 6.1 Sol!&quot; and so I clicked the button to try it, and intriguingly, it set my model selector to &quot;GPT-6 Astra Light&quot; which makes me think 6.1 Sol may be in some way just a lighter&#x2F;distilled version of Astra? defo interesting, not sure if I should read too much into it though. I see now option for directly selecting 6.1 Sol in my Codex Desktop UI.
    1. slekker · · focus · HN ↗
      Astra Light is the default option in the UI, so likely a bug
      1. recursive · · focus · HN ↗
        Wait, so is coding not solved?
        1. alirezaxdehghan · · focus · HN ↗
          more like testing is not solved
      2. skerit · · focus · HN ↗
        So their original plan was to axe Terra, but then introduce an &quot;Astra Light&quot; model a week later? They had a nice lineup named for a whole 3 months, and they&#x27;re already messing with it.
    2. solarkraft · · focus · HN ↗
      I wouldn&#x27;t read too much into it. I&#x27;ve gotten this popup for a model I hadn&#x27;t had access to yet before and it resulted in what you describe.
    3. netruk44 · · focus · HN ↗
      [delayed]
  44. lynx97 · · focus · HN ↗
    Came here only to check if the pelican spam has made it to the top again.
  45. Aboutplants · · focus · HN ↗
    “OpenAI&#x27;s new Pro 500 plan offers OpenAI&#x27;s highest usage allowance and comes with access to its new &quot;Ultrafast&quot; feature — it also costs $500 per month.

    At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”

    <a href="https:&#x2F;&#x2F;www.engadget.com&#x2F;2272106&#x2F;openai-adds-dollar500-pro-subscription-nerfs-its-existing-dollar200-tier&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.engadget.com&#x2F;2272106&#x2F;openai-adds-dollar500-pro-s...

    Yikes

    1. surgical_fire · · focus · HN ↗
      OpenAI is deeply unprofitable, particularly on those pro plans.

      The only way is for prices to go up. Way up.

      1. mrtesthah · · focus · HN ↗
        It really does look like OpenAI is trying to gradually get rid of their subscription plans. Every week there is noticeably less usage available to them while each new model release boasts substantially cheaper API token pricing. If this continues then the two pricing models will eventually be at parity.
        1. surgical_fire · · focus · HN ↗
          The subscription plans are a huge money sink, that obviously will have to go away or be priced at a ridiculous level to make sense.
          1. machomaster · · focus · HN ↗
            This is not true at all, at the most fundamental level. There is a reason why all the businesses (IT, gyms, cars, restaurants, streaming services, music, games, stores, apps, food delivery, magazines, newspapers, shaving blades, parfume, etc) are doing everything in their to get subscribers and are willing to decrease prices in order to get customers who are paying the monthly (or even better, a yearly) fee.
            1. surgical_fire · · focus · HN ↗
              This is delusional.

              The cost of providing the tokens for a heavy user (and let&#x27;s be frank, the people paying $200 are likely heavy users) is many, many times more than the $200 recurring revenue they generate.

              1. machomaster · · focus · HN ↗
                Let&#x27;s be frank, you don&#x27;t know what you are talking about.

                Deepseek has low prices and despite that their profit margin at the beginning of this year was a whooping 82.9%. Since then, they have significantly raised prices.

                You can actually check the approx. financials of OpenAI and Anthropic. The growth is insane.

                There is no reason to believe why OAI&#x2F;Anthropic wouldn&#x27;t have a much better profit margin than DS, taking into account a much higher prices.

                1. surgical_fire · · focus · HN ↗
                  rofl, and I am the one that doesn&#x27;t know what he is talking about.

                  &gt; Deepseek has low prices and despite that their profit margin at the beginning of this year was a whooping 82.9%. Since then, they have significantly raised prices.

                  DeepSeek increased prices substantially not long ago. I find their profit margins hard to inspect considering I have very little idea what sort of environment they may get in China (from cheaper energy to government subsidies). I honestly doubt you have any insight here as well.

                  &gt; You can actually check the approx. financials of OpenAI and Anthropic.

                  No you can&#x27;t. They are not publicly traded, and they constantly and selectively leak bullshit metrics, from extremely unclear ARR, to extremely deceiving EBITDA. You willingly eat their bullshit and call me a picky eater in return.

                  &gt; There is no reason to believe why OAI&#x2F;Anthropic wouldn&#x27;t have a much better profit margin than DS, taking into account a much higher prices.

                  I see no reason to believe (much less any actual evidence) that OAI or Anthropic have any path to profitability.

                  If inference (particularly for subscriptions) was in anyway as profitable as you claim today, they wouldn&#x27;t need private investment rounds like crazy nor they would be desperate to offload this hot potato in an IPO.

                  82% margins lol. Are you telling me that if you created a machine that turns 1 dollar in 5 what you would do is dillute your ownership of the machine instead of using these fabulous profits to expand the business?

                  1. machomaster · · focus · HN ↗
                    Further proof of my previous verdict...

                    It&#x27;s clear that you are out of your depth when it comes to financials, business economics or a simple &quot;what it takes to run a business&quot;.

                    You need money to make money. Growth strategy vs. self-financing strategy, pros and cons, when to do each. Critical chain. Limiting factor in infrastructure. Will not expand because this already goes over your head.

                    1. surgical_fire · · focus · HN ↗
                      Yes, and apparently they need several trillions of revenue to make the investments make any sense.

                      I&#x27;m not the one out of my depth here.

                      Feel free to have the last word. I prefer to read idiocy in homeopatic doses.

    2. moregrist · · focus · HN ↗
      This is pretty typical product positioning. You want to sell to both high-end and low-end users, so you offer products at a few price points. Then it turns out that that middle is a much better fit for most users. So you start making the middle a worse fit to push most of those users into the higher tiers.

      Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We&#x27;ll eventually learn whether both are true. For OpenAI right now, it&#x27;s probably enough to just increase revenue, even if the higher tier is even less profitable.

      1. 5555watch · · focus · HN ↗
        The 200$ plan was appealing because you got 4x usage for 2x the price.

        Now, as it&#x27;s linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.

        With this in mind, it sounds like a fumble by OAI.

        1. jpadkins · · focus · HN ↗
          This is what I did. Hope it works out. The other benefit is you have a more natural method to avoid lock in. A lot of &quot;improvements&quot; to the agent harness I believe are attempts to build customer lock in.
        2. RussianCow · · focus · HN ↗
          &gt; it sounds like a fumble by OAI.

          The vast majority of their revenue comes from large businesses buying for their teams, which are almost certainly not going to juggle lower tiers of different subscriptions to save a few bucks.

          1. nananana9 · · focus · HN ↗
            If we&#x27;re heading to a world where AI spending for companies will be close to salary spending - which is questionable, but is the only way future in which OAI&#x2F;Anthropic survive - you will most certainly have people whose full-time job it is to juggle providers and figure out how to save a few percent this month.
          2. 5555watch · · focus · HN ↗
            Maybe. But developers are also private people who will also play with their private subs and projects. In my opinion, their private experience might influence some corporate level decisions.

            Being grandfathered by OAI and happy is not the same as having both, and noticing &quot;hmm maybe Claude is much better for my case, Ill suggest that to our manager&quot;

            1. RussianCow · · focus · HN ↗
              [delayed]
    3. TomGarden · · focus · HN ↗
      They&#x27;re really (finally?) starting to behave like a company bleeding money.

      Our VC-backed subscription days are numbered

      1. Forgeties79 · · focus · HN ↗
        They spent 48bill and made like 4bill last year. I can’t imagine things have gotten much better over there. The squeeze is definitely coming.
      2. m3kw9 · · focus · HN ↗
        I&#x27;m ok with whatever price they give out given they are not a monopoly and have competition, the lock in is minimum for me. This means they have legit reasons to send us this price plan. I don&#x27;t believe they would shoot themselves in the foot when there is cut throat competition (Claude&#x2F;opensource) out there.

        Lastly, I&#x27;d like to actually use it in the real world to see how far my plan goes or if its unusable.

      3. glaslong · · focus · HN ↗
        Alas, I did enjoy burning investor money on my taxis, movies and tokens.
      4. onlyrealcuzzo · · focus · HN ↗
        &gt; Our VC-backed subscription days are numbered

        Well, the time it takes to compress frontier intelligence down to DeepSeek V4.1 Flash costs (basically too cheap to meter) is dropping, and the differential between the two is also dropping...

        So... who cares?

    4. honkycat · · focus · HN ↗
      Wow, canceling my sub. Lets see how Claude is doing these days.

      I can justify $200&#x2F;mo but more than double is not appealing to me.

      1. WinstonSmith84 · · focus · HN ↗
        Well, here is a breaking-news for you: the 20x from Claude is not a 20x on the weekly usage, it&#x27;s a 20x on the 5h usage, while the weekly usage is simply double the $100 plan...

        Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn&#x27;t have a 5h limit.

        1. diffuse_l · · focus · HN ↗
          OpenAI 20x wasn&#x27;t 20x even before that change. I got a lot more from Claude 5x than Codex 20x...
          1. spiderice · · focus · HN ↗
            You are literally completely flipping reality. Codex was, in fact, 20x. It was Claude that was not 20x until they got caught.
            1. diffuse_l · · focus · HN ↗
              I&#x27;m describing what I got from 20x Codex vs Claude 5x. Codex is just not worth the money, at least for me. What&#x27;s flipped is the value you get for each of those
        2. cmrdporcupine · · focus · HN ↗
          &quot;For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. &quot; - Dario a couple weeks ago.

          Yes, he was talking about safety, but IMHO they&#x27;re likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.

          I suspect we&#x27;ll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.

        3. enraged_camel · · focus · HN ↗
          You are painting half of the picture, perhaps on purpose? The other half is this: Opus 5.5 is significantly better than both Sol 6.1 and Astra, and with the newly increased limits across the board, it is quite difficult to run out (unless you&#x27;re spamming agents at Max effort). So it is a much, much better deal than OpenAI&#x27;s Pro 100.
          1. WinstonSmith84 · · focus · HN ↗
            &gt; Opus 5.5 is significantly better than (..) Sol 6.1

            Come on .. this is barely released and you can already make that assessment?

            And no, the $200 Anthropic plan is not significantly better than the $200 OpenAI plan, it&#x27;s just the same Marketing non-sense and anybody shall now rather stick to the $100 plan of both of these provider if the monthly budget is $200. Anthropic doesn&#x27;t have a Luna Max equivalent, and frankly Sol 6.1 is yet to be thoroughly tested.

          2. [deleted] · · focus · HN ↗

            [deleted]

        4. MCArth · · focus · HN ↗
          If you&#x27;ve used both you know the OpenAI plans don&#x27;t compare to Anthropic plans _at all_. Claude code subscriptions are probably worth 4x as much in API spend compared to the same OpenAI subscription tier.
          1. nostrebored · · focus · HN ↗
            I think you have probably started using OpenAI recently -- one draw used to be that it was really, really hard to ever hit limits. If you did, you probably had usage resets available.

            I think this is still true provided you&#x27;re not using Astra.

            1. machomaster · · focus · HN ↗
              The shitty thing about OpenAI&#x27;s resets is that, unlike Anthropic, they also reset the limit (on the next natural weekly reset). It means that of you pushed the reset button 5 days into the week, you only get 2 days&#x27; (2&#x2F;7 of weekly) worth of extra tokens.
              1. cromka · · focus · HN ↗
                Most recent two resets were banked?
                1. machomaster · · focus · HN ↗
                  Yes, as I mentioned as well.

                  An example. Let&#x27;s assume that the work is evenly divided between days.

                  Imagine you want to work twice as much.

                  1. How efficiently can you use the reset credits if they would not reset the normal reset time?

                  Work with your normal weekly quota 3.5 days, press reset, work with new tokens for the rest of the week. Efficiency 100%.

                  2. With natural reset time going forward 7 days after each artificial reset.

                  You work for 3.5 days, press the reset, work for 3.5 days, wait for another 3.5 days for the natural reset, work for 3.5 days, press manual reset, work for 3.5 days, wait for 3.5 days... You can calculate the number for decreased efficiency yourself.

                  1. cromka · · focus · HN ↗
                    Oh right, I missed your point. Yeah I guess those resets are only for when your workload required you to use all the allowance before end of week. Then you reset and are back to fresh &quot;regular&quot; (non-overloaded) week. But I agree it would be nicer if it worked like you wish.
          2. andriy_koval · · focus · HN ↗
            people say this, but I am wondering if there is benchmark&#x2F;dashboard which actually measures this?
        5. the_duke · · focus · HN ↗
          It used to be bad, but right now with the 200$ Claude sub I find it pretty hard to blow past the session limit.

          You have to do a lot of things in parallel.

          1. InsideOutSanta · · focus · HN ↗
            Yeah, Fable is essentially unusable, it just burns through quota, but Opus 5.5 is great. The $200 plan goes a long way.
      2. spiderice · · focus · HN ↗
        &gt; According to the company, existing subscribers will keep their current limits for a time, and will later receive a one-time credit to help them make the most of their new reduced allowances

        Might want to hold off on canceling and continue to bleed them dry until the nerf hits

        1. honkycat · · focus · HN ↗
          I just got an email telling me this isn&#x27;t true. They&#x27;re immediately cutting my 200, which I&#x27;ve had for like a year.
    5. torginus · · focus · HN ↗
      I think that&#x27;s by design - they&#x27;re going to IPO soon so if they can get a significant percentage of users to switch from the $200 to the $500, they can 2.5x projected revenue.
      1. adonese · · focus · HN ↗
        Very risky to do so especially considering how well is opus 5.5.
        1. scottLobster · · focus · HN ↗
          You think these guys care about risk?
          1. cromka · · focus · HN ↗
            Their investors do
      2. glub · · focus · HN ↗
        Yeah, that&#x27;s not going to happen. They are more likely to lose a lot of customers, unless Anthropic does the same thing.

        But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.

        1. latentsea · · focus · HN ↗
          For consumers they may as well buy GPUs and run local models. The cost is same over a year or two but infinite token usage, they get to keep the hardware, and local models continue to improve over that time too. I can&#x27;t justify $200 on SOTA models for a personal subscription after Qwen3.8-27B. And it&#x27;s only getting better from here.
          1. glub · · focus · HN ↗
            Yes, either US AI corps reduce the cost of their top tier personal subscriptions down to what people are already paying for other expensive personal apps (e.g. Adobe), so ~$50-100, or open weights are going to eat their lunch very quickly. We&#x27;re not there yet, as current hardware doesn&#x27;t allow you to do things like multiple parallel agents, but we&#x27;ll get there soon enough.

            $500 for the old $200 is definitely a fumble.

            1. latentsea · · focus · HN ↗
              I have multiple GPUs now as a way to solve that.
              1. rrvsh · · focus · HN ↗
                Surely you realize how rare the ability to do this is
                1. latentsea · · focus · HN ↗
                  I do not. I had 1 GPU and I had an expensive subscription. I simply cancelled it, and purchased a second figuring if I was going to spend the money anyway I&#x27;d rather have something to show for it at the end of the day. I&#x27;m not unique or special in my capability to do this. I figure I may as well purchase at least one GPU per year equivalent to what I would have spent on SOTA model subscriptions for that given year.
          2. RussianCow · · focus · HN ↗
            People keep saying this but it&#x27;s just patently not true, or at least not apples-to-apples. You can&#x27;t seriously compare Qwen 3.8 27B to Fable or Astra. Even if local models get better, so will the frontier, and you&#x27;ll always be at a disadvantage.

            Unless you&#x27;re talking about buying enough hardware to run something like GLM 5.3, in which case the math just doesn&#x27;t pencil out—the break even point is several years, and you&#x27;re stuck with hardware that will be outdated well before then.

            There are plenty of good reasons to use local models, but none of them are financial, at least for the vast majority of users.

            1. latentsea · · focus · HN ↗
              You don&#x27;t need SOTA. You need a model that can accomplish your task. Qwen3.8-27B isn&#x27;t comparable to SOTA, but can I use it and accomplish most of my tasks with? Yup.

              The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.

              This way you&#x27;re not actually at any disadvantage in terms of capability. You also don&#x27;t need an advantage, you need to complete the tasks you care about. Eyes on the prize.

              RTX 3090 came out a long time ago and it may be &#x27;outdated&#x27; at this point but still banging like a champ for anyone who bought one and becoming increasingly more capable as new models unlock it&#x27;s potential. Hardware hasn&#x27;t changed much, but what it can do certainly has.

              1. RussianCow · · focus · HN ↗
                [delayed]
        2. seizethecheese · · focus · HN ↗
          [delayed]
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.