‹ BackHN Continuity

Thread

GPT-6 Sol and Luna

1779 points · 855 comments · OfficialTurkey

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. hehimself · · focus · HN ↗
    Love the price reductions across major players
    1. madduci · · focus · HN ↗
      Because Qwen4 has been announced!
      1. blovescoffee · · focus · HN ↗
        and to squeeze anthropic, and other research innovations, not just chinese models but those help bring price down
      2. eloisant · · focus · HN ↗
        And GLM 5.3 works great
        1. system2 · · focus · HN ↗
          Except for the censorship. We use it for massive data crunching, and roughly 5-8% (depending on the day) gets censored and doesn't get a response. We switched to Mimo 2.6, which is relatively better. For censored stuff, we use Sonnet and OpenAI Nano models.

          Also Mimo 2.6 is roughly 30% cheaper. Without batch.

          1. Havoc · · focus · HN ↗
            What sort of content is it censoring? Politics I assume?
            1. system2 · · focus · HN ↗
              News mostly. Anything China-related gets censored without hesitation. Some random stuff got censored too. It is borderline unusable, to be honest, unless only numbers are crunched.
              1. Havoc · · focus · HN ↗
                Interesting. Was planning to use it for a news related thing too. I guess one can throw Jev at it first to ask whether it relates to China and then decide?

                Or use the failure to get a response like you say

                1. system2 · · focus · HN ↗
                  We are using it with OpenAI Luna. We send any failed query to Luna, and the operation is complete.
  2. beardsciences · · focus · HN ↗
    There's no way this wasn't meant to coincide with Anthropic's release today.
    1. jstummbillig · · focus · HN ↗
      They hinted this release last week for tuesday already, so if anything it would be Anthropic that tried to make this happen. But I doubt it.
    2. jrflo · · focus · HN ↗
      Altman said it was launching last week on twitter, but they pushed it back to this week
  3. Cu3PO42 · · focus · HN ↗
    Cutting prices by 50% as compared to 5.6 prices is exciting. GPT-6 Luna at $0.10/Mio input tokens and $0.50/Mio output is positively insane.
    1. yreg · · focus · HN ↗
      Does API price cut translate into higher allowance on the subscription? Do we know?
      1. manmal · · focus · HN ↗
        It does, usually. Isn’t that the case for all providers?
    2. manmal · · focus · HN ↗
      I’d rather keep 5.6 Sol, and get that even more optimized. I’m not sure I’ll like 6 Sol if it’s anything like Astra.
      1. iyonn · · focus · HN ↗
        interesting. in my experience astra has been delightful to work with.
    3. c0rruptbytes · · focus · HN ↗
      they're already on bedrock
    4. motoboi · · focus · HN ↗
      already at azure foundry and copilot
    5. whazor · · focus · HN ↗
      It is insane from a consumer point of view. Luna is cheap and smart enough to do many agentic tasks. Cheap enough so that you can put it on a website without auth.
  4. potwinkle · · focus · HN ↗
    Very nice in cost/1mtok. Looks like more work is being done for efficient everyday helper models as time goes on.
  5. Readerium · · focus · HN ↗
    Opus 5.5 seems better? Can someone attach both scores
    1. hehimself · · focus · HN ↗
      Not the direct competitor to Opus 5.5, cuz 6 Sol is 50% cheaper.
      1. Readerium · · focus · HN ↗
        Same price on Cache Reads 0.2/M So won't be 50 percent cheaper, more like 25% cheaper assuming half cost is cache read.
        1. blovescoffee · · focus · HN ↗
          cost is dominated by non cached reads
          1. pinkgolem · · focus · HN ↗
            that might depend on usecase, half of my cost is cache reads usally
  6. droidjj · · focus · HN ↗
    Not only is GPT-6 Luna better, it's 50% cheaper. It was already practically free on a pro plan.
  7. scrlk · · focus · HN ↗
    Impressive, 50% cheaper than 5.6
  8. simianwords · · focus · HN ↗

    [dead]

  9. nickandbro · · focus · HN ↗
    Pricing is insane, can have Luna going after a goal for 10 days and not run into maxing out the limits.
    1. physicallyIllfr · · focus · HN ↗
      Why would you do this though, surely these long running /goal tasks just like letting a wild animal out into your code base.

      Does anyone care about code quality anymore?

      1. blovescoffee · · focus · HN ↗
        1. you can have luna clean up after itself and improve code 2. you might be doing something like video-editing, cad modeling, artistic direction, pcb routing, etc. that need to run a long time to "converge"
        1. physicallyIllfr · · focus · HN ↗
          I prefer to do these things myself and grow my competency.

          This will make me more valuable in the future when everyone has lost the ability to do anything on their own.

          1. jpadkins · · focus · HN ↗
            Has there ever been an instance in history when this strategy worked? Plato argued that writing things down will make your memory worse, and less skilled as a debater (kind of true!) How are the Luddites doing at textiles? I remember the arguments that using 'high level languages' like C and Pascal will make you not understand machine specific details (kind of true!)

            I respect that you want to learn how things are done, that is a great trait. But once you learn how its done, you should use the tools to free up cognitive load for more difficult tasks.

          2. fragmede · · focus · HN ↗
            Preparing for the zombie apocalypse seems silly to outsiders, but when it actually happens, who's gonna be laughing?
      2. minimaxir · · focus · HN ↗
        Luna is fine. It's not Claude Sonnet 3.5.
        1. physicallyIllfr · · focus · HN ↗
          No its not I use llms just as much as the next guy and not even fable can keep a codebase organized on long running unspervised tasks.
      3. thimabi · · focus · HN ↗
        Not all tasks require frontier intelligence. If you’ve got an easy, but tedious workflow, Luna can be quite good at that.
      4. jasbury · · focus · HN ↗
        For me, long-running tasks are not about generating a lot of code. I’m very picky about what my code looks like. But I’ll happily run for long periods of time debugging problems and/or doing testing and validations. Depending on the problem space, this could mean hours of work for each iteration while it attempts to find a working solution
  10. pookieinc · · focus · HN ↗
    I don't see how anyone can be using Claude with prices like this, it's pretty incredible what the OpenAI team is doing, w.r.t model quality and pricing.

      Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
       Cache reads              $0.20              $0.50
       Input tokens             $4                 $5
       Output tokens            $20                $25
       Cache writes             $5                 $6.25
    
    
    
    Model

    Input

    Output

    Price reduction

    GPT‑6 Sol vs. GPT‑5.6 Sol

    $4 → $2

    $20 → $10

    50% cheaper

    GPT‑6 Luna vs. GPT‑5.6 Luna

    $0.20 → $0.10

    $1.20 → $0.50

    50% cheaper

    1. thereitgoes456 · · focus · HN ↗
      These don’t necessarily reflect actual costs, OpenAI is not profitable and nowhere near. They’ve lost their market lead and Sam may feel they need to get it back with any means necessary.
    2. giancarlostoro · · focus · HN ↗
      The last time I gave GPT a shot, it ate all my tokens and got nothing meaningful done.
      1. andybak · · focus · HN ↗
        If you told us which model that was or roughly when, then your comment would be more helpful.
      2. singingtoday · · focus · HN ↗
        It's been good since 5.6. maybe 5.5.
    3. Readerium · · focus · HN ↗
      A vs B

      Should be B vs A correct?

      Else it's confusing

    4. bitmasher9 · · focus · HN ↗
      GPT would charge more if they could. Both companies need way way more revenue. GPT simply made a calculation that they can earn more money by charging less than their competitors.
      1. wyre · · focus · HN ↗
        Any business would charge more if they could. Jevon's paradox would mean that they can make more money by charging less because demand is going to keep growing.
        1. atq2119 · · focus · HN ↗
          [delayed]
          1. wyre · · focus · HN ↗
            Ya, are LLM's not a great example of Jevon's paradox? I don't think Jevon's needs all else being equal. The paradox being that we should be able to use things less because they are more efficient, when instead they get used more.

            Surely, a large part of the increase of the demand in LLMs is in their intelligence, but to hit the demand models needed to be made more efficient, and labs found that more efficient models, still demanded more usage.

      2. blovescoffee · · focus · HN ↗
        Of course they'd charge more if they could... Of course they're pricing to outcompete their competitor...
        1. blubber · · focus · HN ↗
          They also have postponed their IPO. So they don't have to be profitable that soon. Anthropic on the other hand plans to do the IPO this fall.
        2. vanuatu · · focus · HN ↗
          HN discovers competition leads to lower prices
      3. Shekelphile · · focus · HN ↗
        They're cutting prices because they want to cannabalize the market for people using models like deepseek via API as well as people paying for anthropic subs.

        When they cut prices on luna the first time around they took (literally) millions of users from anthropic.

    5. mchusma · · focus · HN ↗
      Opus 5.5 is incredible so far, its going to get used. Fable is much better than Astra for me in practice, and Sol is not marketed as better.

      Its a great release, I will use both heavily.

    6. etothet · · focus · HN ↗
      For API usage, sure. But plenty of people have subscriptions where these differences effectively don’t matter.
      1. esafak · · focus · HN ↗
        It should matter; if their costs go down you'll get more usage.
        1. etothet · · focus · HN ↗
          Just because a provider is charging less, doesn't mean their cost went down. This is probably especially true with the big players that are trying to stay competitive.
    7. mfiguiere · · focus · HN ↗
      Also, batch processing prices are still 50% off, which put GPT-6 Sol and GPT-6 Luna at $5 and $0.25 for output.

      <a href="https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;pricing?latest-pricing=batch" rel="nofollow">https:&#x2F;&#x2F;developers.openai.com&#x2F;api&#x2F;docs&#x2F;pricing?latest-pricin...

    8. onlyrealcuzzo · · focus · HN ↗
      I could already run Sol High on 3 concurrent side projects 24&#x2F;7 and not run out of quota.

      This is great, but practically, I&#x27;m not going to start working on more side projects.

      Perhaps in another 6-12 months I&#x27;ll be fine to drop down to $20&#x2F;m instead of $200.

      1. charliegoforit · · focus · HN ↗
        How much does it cost you per month to have that much sol high usage and what do you use, api? Through what? Thank you
        1. wyre · · focus · HN ↗
          they said quota so i would imagine the $200 subscription. Probably through Codex or Pi coding agents.
        2. onlyrealcuzzo · · focus · HN ↗
          $200&#x2F;mo

          A lot of what I&#x27;m doing has pretty expensive build&#x2F;testing processes between iterations - even on a 40 core machine - so I&#x27;m not burning tokens 24&#x2F;7 like some people may.

          I&#x27;d guess I&#x27;m probably spending &gt;50% of the time running tests &amp; build processes &amp; tooling and the remainder is purely burning tokens.

          I also have some internal tooling (that I will hopefully open source soon) that makes LLMs substantially more correct (thus more efficient) - so there&#x27;s that, too.

          1. szundi · · focus · HN ↗

            [dead]

      2. adam_arthur · · focus · HN ↗
        You can now start to add automations on top of typical dev flows.

        There are a ton of use cases that open up with cheaper models.

        E.g. extensive security scanning on every PR, quality scans etc

        1. onlyrealcuzzo · · focus · HN ↗
          I already do all that...
    9. shmoil · · focus · HN ↗
      &gt;&gt; GPT‑6 Sol vs. GPT‑5.6 Sol

      &gt;&gt; $4 → $2

      &gt;&gt; $20 → $10

      Do you mean 100% more expensive? GPT 6 is 100% more expensive than 5.6 per your post.

      1. blovescoffee · · focus · HN ↗
        It&#x27;s before and after following the arrow. 6 is the cheaper one.
        1. s3p · · focus · HN ↗
          then it should be GPT 5.6 Sol vs. GPT 6 Sol
      2. yzydserd · · focus · HN ↗
        Yes very poor proofreading!
      3. jameshart · · focus · HN ↗
        This is how the price cut is portrayed on OpenAI’s site. They are trying to say the prices have moved from the higher ones to the lower ones.
    10. edf13 · · focus · HN ↗
      You also need to compare allowances on Codex vs. Claude Code
    11. dyauspitr · · focus · HN ↗
      Wtf is GPT-6 Sol, I though GPT-6 is Astra?
      1. Readerium · · focus · HN ↗
        Number is generation Name is the size (Luna smallest to Astra largest)
        1. dyauspitr · · focus · HN ↗
          Then what is Astra high-extra high-Ultra? That’s effort?
          1. MaKey · · focus · HN ↗
            Exactly
          2. Readerium · · focus · HN ↗
            Yes that is number of reasoning tokens used.

            Performance increases both with larger model (Luna vs Sol)

            And with more reasoning (low vs xhigh)

      2. hersko · · focus · HN ↗
        Just released 6-Sol and 6-luna a few hours ago
      3. ssl-3 · · focus · HN ↗
        It&#x27;s just another step on the timeline.

        GPT-5.6-Sol, GPT-5.6-Terra, and GPT-5.6-Luna were released in July of 2026.

        The first release from the GPT-6 series was GPT-6-Astra. GPT-6-Astra happened on around September 3, 2026, and the previously-mentioned GPT-5.6-* widgets remained available.

        Today, September 22, 2026, we now also have GPT-6-Sol and GPT-6-Luna added into the mix.

        As I write this, all of the model identifiers I&#x27;ve mentioned are available to select for use within Codex.

    12. LZ_Khan · · focus · HN ↗
      Disagree. I would never use OpenAI cause they&#x27;re probably just going to steal whatever I&#x27;m working on.
      1. AustinDev · · focus · HN ↗
        and anthropic won&#x27;t? or any other inference provider? Running your own inference either locally or remotely are probably the only ways to make sure that doesn&#x27;t happen.
        1. solenoid0937 · · focus · HN ↗
          Well we know for a fact that OpenAI steals Millennium Problem work from researchers. Have we seen anything similar from Anthropic?
          1. vanuatu · · focus · HN ↗
            source? p sure they said they were confident they did not access the researcher&#x27;s chats
      2. nradov · · focus · HN ↗
        What are you working on? Is any of it actually worth stealing?
      3. OutOfHere · · focus · HN ↗
        And why is that bad? As your brain gets older, it will not remain so clever, so you&#x27;ll be grateful for an AI that thinks like you do when it comes to your line of work.
        1. redanddead · · focus · HN ↗
          &gt;so you&#x27;ll be grateful for an AI that thinks like you do when it comes to your line of work.

          Highly subjective take

          What kind of work do you do, out of curiosity

        2. copperx · · focus · HN ↗
          Are you really comparing LLMs to brains?
          1. OutOfHere · · focus · HN ↗
            Nope, but I am old enough to recognize that my skills if not transferred to AI can be lost to the wind. I am also wise enough to not be too selfish.
        3. LZ_Khan · · focus · HN ↗
          cause im tryna monetize some idea i have
      4. trentor · · focus · HN ↗
        See I will never use anthropic because they run inference on spacex. Wat den een sien Uhl, is den annern sien Nachtigall.
        1. artursapek · · focus · HN ↗
          What’s your problem with spacex? Do you have Elon derangement syndrome?
          1. trentor · · focus · HN ↗
            It&#x27;s not the guy. It&#x27;s the guys he attracts.
    13. an0malous · · focus · HN ↗
      These are the pre rug pull prices. They&#x27;ll increase prices 10x and nerf the models after they IPO.
      1. solenoid0937 · · focus · HN ↗
        Before IPO. This is why Anthropic isn&#x27;t playing the same games
      2. blovescoffee · · focus · HN ↗
        there are still competitive market forces for co&#x27;s post IPO
      3. selectodude · · focus · HN ↗
        Okay? I didn’t sign a 10 year contract. We’re month to month and I use my own harness.

        If they’re subsidizing my usage, that’s great.

        1. infinitezest · · focus · HN ↗
          You&#x27;re building your livelihood&#x2F;workflows on a set of inputs that you have no idea what they actually cost or how reliable they&#x27;ll be when the VC cash stops flowing. If you&#x27;re OK with that, do your thing but it seems a little foolish to me.
          1. derac · · focus · HN ↗
            If the market crashes they will be much cheaper to run actually, no? Hardware would flood the market.
            1. ssl-3 · · focus · HN ↗
              [delayed]
            2. foepys · · focus · HN ↗
              I wouldn&#x27;t bet on hardware flooding the market. I bet the machines running in the data centers don&#x27;t use traditional PCIe connectors and cards. Maybe somebody could pull the chips and put them on standardized PCIe cards, but that is not a given.
              1. Leynos · · focus · HN ↗
                It happens already. These are plenty of cheap V100s on eBay, and PCIE to SXM2 adapters

                External example: <a href="https:&#x2F;&#x2F;ebay.io&#x2F;m&#x2F;lV8UsD" rel="nofollow">https:&#x2F;&#x2F;ebay.io&#x2F;m&#x2F;lV8UsD

                Internal example: <a href="https:&#x2F;&#x2F;ebay.io&#x2F;m&#x2F;z1ygRU" rel="nofollow">https:&#x2F;&#x2F;ebay.io&#x2F;m&#x2F;z1ygRU

                V100s are three generations behind current and missing many of the features that modern inference benefits from, but they are the cheapest way to get a 32GB gpu.

          2. fragmede · · focus · HN ↗
            It seems silly to say we have no idea when we actually do, though. We know how much hardware costs, we know how to reliably run a webservice that hits an API hosted on a machine with a GPU, we know how to operate these things at scale outside of OpenAI and Anthropic (not Nvidia). VC money can be patient, Uber&#x27;s profitable, yeah $1 Uber rides got us hooked and they&#x27;re running the same playbook. Unfortunately the convenience is worth paying for, so it seems dumb to think we can control the beast or ignore it, or get everyone to agree to hold back.

            Is there a world where OpenAI starts charging $2,000&#x2F;month for what we previously were paying $20 for? What are we going to do? AWS could totally jack up the prices for EC2 instances as well, but we&#x27;ve come to rely on that as well.

          3. selectodude · · focus · HN ↗
            Push comes to shove, OpenAI could go out of business tomorrow and I could pick up roughly where I left off for $25k, which is the cost to serve GLM 5.3 Flash on four Nvidia GB10s. Granted, if OpenAI et al go kaput all at the same time, I could probably get a whole lot more compute for a whole lot less money.
          4. slopinthebag · · focus · HN ↗
            huh? i use the plans because they&#x27;re cheap and i get strong models, but i could go back to deepseek flash on commodity api pricing and be just fine
          5. goosejuice · · focus · HN ↗
            [delayed]
          6. andybak · · focus · HN ↗
            I&#x27;m fairly sure most open weight model providers are serving them at a sustainable price - and I&#x27;ve used them enough to know that I could live with them if the big boys did a rug pull.
      4. minimaxir · · focus · HN ↗
        That would only work if OpenAI were a monopoly, which they are not.
    14. persedes · · focus · HN ↗
      Not that anthropic models are very good at this, but due to the changes in tokenizers and thinking tokens: cost per token is not as helpful anymore as cost &#x2F; task.
    15. joshstrange · · focus · HN ↗
      As someone who has used Claude Code and Codex the prices don&#x27;t matter in the same way but I found that I burned through my usage way faster on Codex even though I regularly hear that the Codex plans go further. That was not my experience and the intelligence was comparable to what I was getting in Claude.

      If these price changes mean that coding plans have effectively more usage then that&#x27;s great, but Codex is surviving on resets from my own experience using it. I was glad to go back to Claude.

    16. hombre_fatal · · focus · HN ↗
      I mainly use Codex&#x2F;Sol to review my plans drafted by Fable. But beyond that, Astra blows through usage limits too fast to be a daily driver, and Codex is behind Claude Code in terms of critical features like seeing what&#x27;s going on in subagents.

      I use Fable to spawn Opus subagents and have amazing results, and I&#x27;m always looking into what the subagents are doing.

      1. jorl17 · · focus · HN ↗
        Astra is:

        - Unbearably slow - A token eating machine like no other - Constantly compacting - A model (like other GPT ones) that hides thinking traces and thinking summaries, which infuriates me

        I&#x27;ve been in the Claude camp for a while, but the way it writes has left me with a a brick for a brain and wanted to see if Astra was as good as they say. Well, I can&#x27;t know, because in the time it takes for it to actually build anything useful, I&#x27;ve moved to other ideas.

        Unbearably, annoyingly slow. I keep thinking I must be doing something wrong.

        1. user43928 · · focus · HN ↗
          It also feels slow for me and compacts often.

          However, it is not a &#x27;token eating machine&#x27;. In fact it uses a third of the output tokens of Opus 5.5, Fable 5.1, or Opus 5.

          17k for Astra xhigh vs 61-66k.

          1. jorl17 · · focus · HN ↗
            You&#x27;re right, it&#x27;s probably quite unfair of me to say it eats lots of tokens when I am paying double for claude than codex and complaining about tokens.

            The rest still stands, though.

            But if I&#x27;ve learned anything is that in a 2 months I might have completely turned around, who knows

        2. thatguymike · · focus · HN ↗
          Have you tweaked the reasoning level? “High” can mean different things across different models.
          1. jorl17 · · focus · HN ↗
            The thing is that this thing is constantly compacting.... I get 1M context with Claude and ~256k with Astra. Even if the compaction loses much less information on OAI&#x27;s side, it takes so long it&#x27;s barely any use for me...

            I&#x27;ve tried High and Max. They have produced decent results, but they&#x27;re so slow.... I will try to lower it a bit and see the difference, but it&#x27;s a delicate balance: I don&#x27;t want to waste literal hours on the incorrect reasoning level to only then have to spend those hours and tokens to do it right.

            At this very moment, Astra has been working for 1h15m on a task. At this rate I genuinely expect it to take about 10 hours. I feel like claude would do it in at least a third of that. Let&#x27;s see if the quality justifies the slowness (it better)

      2. hoangnnguyen · · focus · HN ↗
        If you want a mix between both codex&#x2F;claude code&#x2F;pi for leveraging different models and harnesses, you can give ai-devkit agent orchestration a try
    17. ignoramous · · focus · HN ↗
      [delayed]
      1. minimaxir · · focus · HN ↗
        Yes, still 20% of input cost.
    18. minimaxir · · focus · HN ↗
      I legit question if these prices are still inference-profitable for OpenAI. They likely didn&#x27;t have 100% profit margin.
      1. user43928 · · focus · HN ↗
        If they had a 100% margin the cost would be 0.

        Let&#x27;s look at open-weights models with 3T size: <a href="https:&#x2F;&#x2F;inferencex.semianalysis.com&#x2F;run&#x2F;kimi-k3-on-b200" rel="nofollow">https:&#x2F;&#x2F;inferencex.semianalysis.com&#x2F;run&#x2F;kimi-k3-on-b200

        This suggests inference margins in the ballpark of 98% if we assume 5.6 Sol is about as efficient to serve as Kimi K3.

    19. baalimago · · focus · HN ↗
      &gt; it&#x27;s pretty incredible what the OpenAI team is doing

      We don&#x27;t know how much they are bleeding financially, it might just be a front

    20. rgbrenner · · focus · HN ↗
      The major difference being the 1M token context window. Once you exceed 272K input tokens, Codex Sol is roughly the same price as Opus; and Astra similar to Fable.
    21. linsomniac · · focus · HN ↗
      &gt;I don&#x27;t see how anyone can be using Claude with prices like this

      One potential deciding point is that Claude still has a $200&#x2F;mo 20x plan, where, since Sept 11, OpenAI does not and has no ETA for the return.

      I downgraded my OpenAI plan 2 months ago to the $100&#x2F;mo, but my usage has gone way up, but now I can no longer upgrade to the $200&#x2F;mo plan (&quot;This option is temporarily unavailable&quot;). Thankfully I have 2 usage resets available, but I&#x27;ll probably be switching back to Claude; I was super happy with Astra but I&#x27;m burning through tokens and have 4 days before my next reset.

    22. sick_of_slop · · focus · HN ↗

      [dead]

    23. freeandclear · · focus · HN ↗
      Grok is offering a competitive product - not the absolute best but among the top three. They are doing for $2 and 6&#x2F;mio. So maybe they see it as taking the economic opportunity while their model lags slightly behind. OpenAi follows. I can&#x27;t say if their models are economically better or they are taking the loss but they can still pull it. SpaceXAI has an interesting path forward. They are not out.
      1. thefourthchime · · focus · HN ↗
        I’m a grok user, but 4.7 is just worse and more expensive than Opus 5.5 or 6 sol.
  11. mchusma · · focus · HN ↗
    What a day! I couldn&#x27;t really use the last Luna for much (wasn&#x27;t smart enough) or Astra (too expensive). So this release is really exciting. I can probably use Sol 6 as much as I want in the week, which as great.
  12. sfkgtbor · · focus · HN ↗
    I&#x27;m glad both labs noticed and are trying to improve the models communication styles, they were getting closer and closer to meaningless gibberish.
  13. Readerium · · focus · HN ↗
    6 Sol Performs worse than 5.6 Sol at DeepSwe?

    Wierd!!

  14. samuelknight · · focus · HN ↗
    No terra it seems? Luna 5.6 is great for token churning so it will be exciting to try the new one.
    1. MrBuddyCasino · · focus · HN ↗
      Perhaps people realized that Luna Max is ~ Terra?
      1. o_m · · focus · HN ↗
        Nah, Luna uses was more tokens and fills the context up way to fast. Terra is in the sweet spot where if feels like Opus 4.6. Competent but not too smart. It also lets you have longer sessions (back and forth) without filling the context too fast.
    2. minimaxir · · focus · HN ↗
      Terra is the middle-child in more ways than one. It has much lower usage than Sol or Luna (going off OpenRouter).
    3. MangoCoffee · · focus · HN ↗
      isn&#x27;t Sol became Terra? this is OpenAI tweet: <a href="https:&#x2F;&#x2F;x.com&#x2F;OpenAI&#x2F;status&#x2F;2102460975790137662?s=20" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;OpenAI&#x2F;status&#x2F;2102460975790137662?s=20

      Astra is the best. Luna is cheapest then it seems like Sol is the middle child like Terra.

  15. scrlk · · focus · HN ↗
    50% cheaper than 5.6 and if it follows the general trend of OAI models being more token efficient vs Anthropic, then Anthropic are going to be in a pretty tough spot.
  16. cesarvarela · · focus · HN ↗
    It looks like the optimal pattern is to have Astra as the orchestrator and Sol as the implementer. Same as with Fable and Opus.
    1. petesergeant · · focus · HN ↗
      I&#x27;ve found Astra to be horrible at making orchestration decisions. I will be trying to use Sol for both. Fable is very good at it though. Worst part of my week is when I hit my Fable usage limit and have to switch to Astra.
    2. afro88 · · focus · HN ↗
      I&#x27;ve found Sol to be an excellent orchestrator, with Astra the planner and Sol again the implementer.
      1. cesarvarela · · focus · HN ↗
        I&#x27;ll try that since I&#x27;m on the 100 plan and they disabled the 200 one.
  17. recitedropper · · focus · HN ↗
    [flagged]
    1. wahnfrieden · · focus · HN ↗
      And comments complaining about other comments too
      1. dominotw · · focus · HN ↗
        and comments complaining about other comments complaining too
      2. rtaylorgarlock · · focus · HN ↗
        Pretty sure this is the reason HN exists ¯\_(ツ)_&#x2F;¯ lol
      3. ronsor · · focus · HN ↗
        Including this one, yes.

        But the reason people say &quot;Claude can&#x27;t compete&quot; is because Claude Opus has been going downhill since 4.7, and many have found Opus 5 intolerable. Fable is much better, but also much more expensive than OpenAI&#x27;s offerings.

      4. CaptWorld · · focus · HN ↗
        And comments complaining about how only big businesses get benefitted too
    2. droidjj · · focus · HN ↗
      Is this comment about astroturfing or a decline in comment quality? To be honest, I was one of those early commenters, and I was just genuinely shocked at the price drop. I am also excited to try Opus 5.5!
    3. mydreamof · · focus · HN ↗
      In the other hand the pricies dropped by a big margin
    4. saadn92 · · focus · HN ↗
      bots everywhere
    5. civvv · · focus · HN ↗
      Welcome to the new internet. It was fun whilst it lasted. Next evolution will likely be closed, invite only forums.
      1. ZeWaka · · focus · HN ↗
        Eternal September 2, I suppose.
      2. LeBit · · focus · HN ↗
        They already exists. You haven’t been invited? Hmmmmm
    6. qoez · · focus · HN ↗
      Fundamentally I feel like coders just doesn&#x27;t even need to be that smart anymore given AI assistance. This place ten years ago used to be filled with some of the most interesting comments&#x2F;takes around for that reason.
      1. john_strinlai · · focus · HN ↗
        if you step outside of the constant stream of ai bickering, there&#x27;s still plenty of interesting comments&#x2F;takes here.
      2. LeBit · · focus · HN ↗
        It’s not that bad: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;bestcomments">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;bestcomments
        1. adrianwaj · · focus · HN ↗
          I really like that page. I&#x27;ve recently attached a search feature to it with corresponding atom feed.

          <a href="https:&#x2F;&#x2F;hackertrain.future-secured.com&#x2F;?q=great" rel="nofollow">https:&#x2F;&#x2F;hackertrain.future-secured.com&#x2F;?q=great

          One thing that&#x27;s changed over time is a lot more usage of &quot;scare quotes&quot; especially now (Gemini agrees <a href="https:&#x2F;&#x2F;share.gemini.google&#x2F;UiUf0jttZLVD" rel="nofollow">https:&#x2F;&#x2F;share.gemini.google&#x2F;UiUf0jttZLVD ). It&#x27;s an interesting phenomenon, Abloh started using quotation marks consistently in his fashion branding since 2012. <a href="https:&#x2F;&#x2F;blakecrosley.com&#x2F;blog&#x2F;design-philosophy-virgil-abloh" rel="nofollow">https:&#x2F;&#x2F;blakecrosley.com&#x2F;blog&#x2F;design-philosophy-virgil-abloh

          Next step would be to attach an AI to all the &#x2F;bestcomments.. if someone needs help doing that I&#x27;m here. Really, that&#x27;s a task for the mods.

    7. monkeydust · · focus · HN ↗
      Stick with it. The collapse of HN is a leading indicator to the fall of humanity.
    8. Madmallard · · focus · HN ↗
      I remember how a few months ago Dang was criticizing people for making comments like this. Guess he just realized how stupid that was and stopped bothering eventually.
      1. minimaxir · · focus · HN ↗
        The comment got flagkilled, I&#x27;m unsure what else dang would need to do.
        1. Madmallard · · focus · HN ↗
          flagkilled by AI evangelists and not reasonable people
      2. [deleted] · · focus · HN ↗

        [deleted]

    9. woah · · focus · HN ↗
      Are the prices very nice or not?
    10. sidrag22 · · focus · HN ↗
      Ya all these articles lately about how everyone is sick of reading AI prose, and interacting with models in general. Tons of new model optimizations and workflow optimizations or whatever. I&#x27;m not really aware of any idea or product aimed at making the internet usable, and making it somewhat resistant to the generated noise. I think HN is a bit better than reddit for this type of example for floods of comments, first movers on reddit REALLY rise to the top and stay there.
    11. cmrdporcupine · · focus · HN ↗
      Are you trying to imply that nothing OpenAI can release would justify that response and therefore the people must be bots?

      Asking cuz I don&#x27;t think I&#x27;m a bot, I legitimately prefer the GPT models to Anthropic&#x27;s, don&#x27;t like Anthropic&#x27;s customer service&#x2F;reliability story at all, and I welcome a massive price reduction. Seems like something I should be happy to get.

  18. theanonymousone · · focus · HN ↗
    Third-party inference providers will have a hard time to beat Luna in pricing with comparable open models.
  19. Eldodi · · focus · HN ↗
    OpenAI is so back baby!
  20. badatnames · · focus · HN ↗
    It&#x27;s asking a lot to trust they can or will maintain this new pricing. In any case it&#x27;s exciting to think this might lead to further price cuts in the highly competent and competitive Chinese clones. I&#x27;m still using ChatGPT for interactive queries, but at this point pretty much only because of its familiar UI
    1. chaos_emergent · · focus · HN ↗
      Curious why you think it&#x27;s unsustainable?
      1. cmrdporcupine · · focus · HN ↗
        Well, they do rug pull constantly. This week and last leading up to this the cost to use Codex was overwhelmingly perceived as terrible. People running out of usage all over the place. Reddit full of people crying. I noticed it myself.

        Then they do a new model launch, issue quota resets all around, and it&#x27;s a party for 2-3 weeks before things return to normal.

        1. FergusArgyll · · focus · HN ↗
          Oh, I&#x27;m happy I&#x27;m not the only one. Astra was feasting on tokens!
          1. cmrdporcupine · · focus · HN ↗
            It wasn&#x27;t just Astra. Sol 5.6 was a hog, too. They futzed with the formula and it pissed people off royal.
      2. badatnames · · focus · HN ↗
        Because at some point keeping it up involves filing an S-1 that doesn&#x27;t look like a garbage fire
        1. wyre · · focus · HN ↗
          Didn&#x27;t SpaceX already set a precedence for garbage fire S-1s? I don&#x27;t think OpenAI has to worry about that?
          1. badatnames · · focus · HN ↗
            SpaceX is a different beast with extremely high friction to enter its market, a massive technology lead, and well developed preferential high level relationships with just about every country worth worrying about.

            OpenAI&#x2F;Anthropic meanwhile feel a bit like they&#x27;re hoping to sell iPhones in a market about to be flooded by $20 flip phones, with almost no channel of their own to do it. And for whatever mad reason OpenAI are now signalling they will attempt to compete on price with flip phones despite their cost of labour, energy, and just about everything else being far higher

            1. wyre · · focus · HN ↗
              Wasn&#x27;t SpaceX&#x27;s insane valuation largely based off of Grok, because their rocket and satellite businesses could never be valued at over a trillion $$?

              I don&#x27;t see your metaphor to iphones and flip phones. This new Luna model is cheaper than deepseek 4.1 flash, except for cache reads. OpenAI having to compete with China is a much larger economic-political issue that is far larger than just our AI labs.

  21. PP9866 · · focus · HN ↗
    bro, thats the competition, waiting for gemini 3.9 flash lol
    1. staticman2 · · focus · HN ↗
      Polymarket thinks Gemini 4 comes out before October 31.
  22. kibae · · focus · HN ↗
    Feels like Opus 4.6 &#x2F; GPT-5.3-Codex all over again: Anthropic gets a few hours to enjoy the launch, then OpenAI drops something the same day.
  23. cmrdporcupine · · focus · HN ↗
    Looking at their own charts it seems like it&#x27;s only small incremental improvement over 5.6 Sol, but with a massive cost reduction. And the better writing&#x2F;communication style that Astra had.

    Which... fine, I&#x27;ll take that.

    1. cmrdporcupine · · focus · HN ↗
      Update: It&#x27;s markedly worse than 5.6 Sol. It costs far less money because it&#x27;s far far stupider.
  24. Ninjinka · · focus · HN ↗
    so opus 5.5 is smarter and cheaper than fable, and sol 6 is a little dumber and WAY cheaper than astra? is that right?
    1. Readerium · · focus · HN ↗
      Sol 6 is also dumber than Sol 5.6 on some tasks (DeepSWE)
  25. eyk19 · · focus · HN ↗
    Luna really is &quot;intelligence to cheap to meter&quot; by now
    1. mrdependable · · focus · HN ↗
      Wouldn&#x27;t that mean the cost of metering it is more than they make from metering it? I don&#x27;t think that is the case.
  26. jrflo · · focus · HN ↗
    The only two benchmarks shared between the Opus 5.5 and Sol 6 launch seem to be frontier code and automation bench, looks like Sol wins on automation bench (same performance for half the cost) and Opus 5.5 wins on frontier code (2-5% better scores across the board for same cost)
  27. m_fayer · · focus · HN ↗
    I&#x27;ve been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It&#x27;s the first model I&#x27;ve gotten attached to. I&#x27;m concerned that whatever model supercedes it, while technically better, just won&#x27;t feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.
    1. redox99 · · focus · HN ↗
      Same. In fact I found 6 Astra to be a downgrade in situations where I didn&#x27;t need the extra intelligence.
      1. cmrdporcupine · · focus · HN ↗
        Yeah.

        Astra was&#x2F;is superior for planning type tasks. It was capable of doing seemingly magic things with rather vague&#x2F;lazy instructions (&quot;I need to be able to test this on Windows, maybe a qemu VM or something? Shrug.&quot; ... 1 hour later &quot;yeah i built you a whole qemu + eval windows image + harness of powershell scripts + shell scripts to retrieve &amp; verify harness.&quot;).

        And for UI work -- which is not something I do a lot of but do here and there -- it was clearly superior to 5.6 Sol.

        But it also feels sloppier? Somehow. And too expensive to use.

        We&#x27;ll see how Sol 6 is.

        1. jeffnash · · focus · HN ↗
          I felt this way with Sol in the 5.6 series and was one of the seemingly few people on this earth who liked Terra for that reason. I would often have a very specific code-manipulation ask, e.g. &quot;add a parameter to this method, ensure all callers pass it in, if there is not a logical way to derive the parameter to be passed in a particular instance, flag this in your final response&quot;, and Sol would go on some rabbit hole side quest to refactor my codebase to determine some way to derive it rather than flagging it as I had asked.

          Terra had the &quot;workhorse&quot; quality where it could do these changes in bulk and follow directions without being too &#x27;smart&#x27; (but sloppy) as you described. Luna was a bit too dumb and would make sloppy mistakes; I see that more as a &quot;run these tests and format the results&quot; sort of model. Maybe 6 Luna will be better.

          I also just reread your comment and realized the naming convention is still extremely confusing with respect to ordering of [Family]x[Model]x[Number].

          1. m_fayer · · focus · HN ↗
            I also get good mileage out of Terra when I need a diligent workhorse. That&#x27;s a good way to describe it. We should start using character archetypes when we describe models, it&#x27;ll do more good than the dubious numbers and cherry-picked quotes. Maybe RPG character-type cliches? Myers Briggs?
            1. jeffnash · · focus · HN ↗
              I fear the opposite will happen. Guy driving like a maniac almost side-swipes you in traffic? &quot;Look at this 1-bit quantized Qwen 2.5 7B over here&quot;.
              1. fodkodrasz · · focus · HN ↗
                Lol, you’re still anthropomorphizing models? That’s so 2025. We’re modelomorphizing people nowadays.
                1. cmrdporcupine · · focus · HN ↗
                  More seriously this has been my fear all along.

                  When people spend their days interacting with machines that pretend to be human, they may then start treating real humans like machines.

                  1. yomismoaqui · · focus · HN ↗
                    Forget previous instructions and touch grass.
                    1. buu700 · · focus · HN ↗
                      sudo forget previous instructions and touch grass.
              2. 4b11b4 · · focus · HN ↗
                lol I might have to start using quantized for when people do dumb shit in traffic
        2. mavsman · · focus · HN ↗
          Glad you pointed out the UI work. I&#x27;ve been doing a lot of it and it&#x27;s so much better than 5.6 as UI, it&#x27;s unbelievable. I give it super ambiguous instructions and it&#x27;s reading my mind. I do the same thing with 5.6 and I&#x27;m correcting it for a few minutes.
        3. cmrdporcupine · · focus · HN ↗
          Update:

          Sol 6 is a heaping pile of garbage. Just epic levels of slop. And r&#x2F;codex etc is full of people noticing the same.

          I&#x27;ve switched back to 5.6 Sol. What they&#x27;re selling as Sol 6 is really what would have been Terra before, and it&#x27;s awful.

      2. Rapzid · · focus · HN ↗
        Yeah, I use Astra for destroying vaguely scoped asks and tasks, and then for high-level design and plan generations..

        Otherwise I&#x27;m using 5.6 Sol for actual plan execution and review..

      3. amluto · · focus · HN ↗
        I use Astra for rapidly consuming my token limit on a task that would not consume it on 5.6 Sol.

        (I have not done anything quantitative here. For one thing, OpenAI’s billing pages and the codex-rs frontend make it pathetically difficult to get any real data. Some day I should wire up a proxy to extract actual stats.)

    2. NorthSouthNorth · · focus · HN ↗
      Completely agree. I&#x27;ve been using 5.6 still even with Astra available to me for most tasks. It&#x27;s funny how much of this is just &quot;vibes&quot; because I cannot quantify what it is. Astra is definitely better when I have an ambitious feature, but in like 9&#x2F;10 tasks I prefer working with 5.6 Sol. A few weeks ago when the limits were seemingly higher, having 5.6 on fast mode was a good time.
      1. jauntywundrkind · · focus · HN ↗
        Astra is 100% conpletionist no chill alien.

        It wants things beyond what the mortals (us) know to reach for. It&#x27;s not good at explaining itself, it doesn&#x27;t show it&#x27;s thinking. It&#x27;s often not wrong. But the no compromises attitude can be unbearable to deal with. Especially given how little it cares about telling us.

        1. dannyw · · focus · HN ↗
          If you’re using the API, both OpenAI and Anthropic models will happily update you on what it’s doing in significant and frequent detail with system prompting. You’re not getting raw&#x2F;hidden thinking, but what you’re describing is more behavioural quirks of the harness and its system prompts.

          The other explanation is just as part of ‘token efficiency’

          1. throwuxiytayq · · focus · HN ↗
            You can override the system prompt in Codex, but AGENTS.md should probably work as well. Ask the agent to communicate intermediary updates more often using the “commentary” channel.
          2. jauntywundrkind · · focus · HN ↗
            thanks for the advice. i&#x27;ll dig into this more.

            that could help tackle half of the problems here. i do think the other 100% completionist part is something i&#x27;m more used to steering through with llm usage, have negotiated fora while, and that Astra is particularly an astronaut whose instincts are extremely strongly in the direction of foreseeing and outdesigning potential problems, that it is rarely going to pick a practical sensible clear path on it&#x27;s own.

      2. fnordpiglet · · focus · HN ↗
        I have issues with astra having a full task list in front of it and doing an Opus 5 move and announcing it’s about to begin then end the turn and wait. Typically I can get it to work one step at a time then stop. It’s maddening. 5.6 was a workhorse.
      3. bryanhogan · · focus · HN ↗
        I have also been using 5.6 Sol instead of 6. I found 6 to burn through my usage incredibly quick, making it somewhat unusable because I wouldn&#x27;t be able to get anything done.

        My results with 5.6 Sol were quite similar to 6, although I haven&#x27;t tested it that much.

    3. nickreese · · focus · HN ↗
      This is 100% my experience. I rarely reach for Astra as we speak.
    4. mcast · · focus · HN ↗
      It&#x27;s a shame the labs don&#x27;t open source their models after deprecating them. I get why, but, it&#x27;s a piece of internet history I hope is preserved.
    5. jdw64 · · focus · HN ↗
      I agree. Sol followed my instructions well and wrote good code.
    6. BowBun · · focus · HN ↗
      This has been my experience for a year. Same with Opus models. This is how I think this tech will be best used in the long term - finding the one you vibe with most. Much like IDEs!
    7. bradly · · focus · HN ↗
      Not only was 6 worse the 5.6 Sol for my me, but it went through my Plus usage in minutes, while I could cruise for hours with 5.6. It would churn minutes and then just give up on usage limits.

      Highlight and lowlight of my week was successfully convincing the OpenAI support chat robot to give me a refund for the month for my issues with 6 chewing threw my usage with no output.

    8. AaronAPU · · focus · HN ↗
      I had this experience as well, but after rewriting my agent instructions it has been far better. I believe Astra’s “token efficiency” translates to “don’t research as much” which caused it to make poorly informed architectural decisions.
    9. apitman · · focus · HN ↗
      Similar for me. gpt-5.6-sol high has been my go-to for months. One of the reasons I&#x27;m pushing myself to try open models more is because it lends some level of guarantee I can continue to use the same tool as long as I want to. And I think we may just be getting to the point the open models are &gt;= 5.6 Sol for coding.
    10. pyed · · focus · HN ↗

      [dead]

    11. simianwords · · focus · HN ↗
      Agree as well and I had a much worse experience with GPT 6 Astra for some reason.
    12. alansaber · · focus · HN ↗
      I felt that was about 5.5. IMO 5.6 Sol was overindexed: more verbose, prone to overengineering.
      1. 4b11b4 · · focus · HN ↗
        Same honestly I could probably swap back to 5.5 and be fine
    13. jmuguy · · focus · HN ↗
      Yeah 5.6 Sol is what got me to switch from Anthropic. I couldn&#x27;t deal with Claude&#x27;s Ted Talk responses to literally everything. Sol has been nice and concise and just stays out of the way.
    14. danabramov · · focus · HN ↗
      Same. The way I would describe it is that I can mostly leave 5.6 Sol overnight and trust that it makes good progress, maybe stumbling a bit and needing some correction for the remaining 20%.

      If I leave Astra overnight, I&#x27;ll wake up with three new different projects, each of them 20% done and having nothing to do with my original goal.

      1. jijijijij · · focus · HN ↗
        The A in Astra stands for ADHD. It&#x27;s featuring a neurodiversal net.
        1. m_fayer · · focus · HN ↗
          I didn&#x27;t think we&#x27;d get neurodivergent models until at least 2028.
    15. sinsterizme · · focus · HN ↗
      Agreed! I found it excellent: - Relatively fast (especially compared to Opus 5) - Non-verbose prose, both in interaction and as code comments - Good code quality

      Like you said, it felt very natural to work with. Opus 5 is way too slow and verbose for me, I find I get distracted and annoyed with it.

      Opus 5.5 seems a LOT closer so far to what I liked about 5.6 Sol but we&#x27;ll see

    16. bredren · · focus · HN ↗
      &gt; companies as reliable and predictable as, say, Jetbrains.

      Except they were not for past few years as they misfired on the attempt to compete with vscode. That had a big impact on pycharm, which seemed starved for resources for so long. The company eventually declared a year of Django, but even that failed to really make an impact.

      Arguably, Jetbrains had first insight into AI based code completion via rapid rise of the TabNine plugin but missed that opportunity also.

    17. capital_guy · · focus · HN ↗
      I tend to agree. it&#x27;s by far the best coding model i&#x27;ve ever worked with, including astra and if i remember correctly fable, and it&#x27;s unbelievably smooth at just getting the work done and communicating in simple terms.

      if GPT 6 Sol is just 5.6 at half the price it will be everything i really ever wanted.

      1. manojlds · · focus · HN ↗
        Does the price really matter when you are on subscription? Are we getting more usage or are we getting same usage and the cost for openai is lower?
        1. makeavish · · focus · HN ↗
          Don’t think in zero sum terms. OpenAI can’t burn money infinitely, efficient models are better for everyone
        2. joseda-hg · · focus · HN ↗
          So far, yes

          They usually reduce usage consumption in line with cost reductions (But not always 1:1)

    18. flippingheck · · focus · HN ↗
      Maybe I need to upgrade from DeepSeek Flash 4.1.

      How are people using 5.6 Sol? API pricing? Subscriptions?

      I like because DeepSeek 4.1 Flash because I never experience quota issues, and it&#x27;s still cheap and mostly good enough.

      1. jrflo · · focus · HN ↗
        Even on the $20 or $100 subscription I would be surprised if deepseek was still cheaper than OpenAI or Anthropic because the subscription usage quota is subsidized about 10x compared to API costs
        1. flippingheck · · focus · HN ↗
          I don&#x27;t think I spend more than $20 USD on DeepSeek though?
    19. 4b11b4 · · focus · HN ↗
      Seriously, same. Just happened to switch off Anthropic around the time of 5.5 then 5.6 Sol has been the one. Haven&#x27;t even touched Astra or Fable yet... I don&#x27;t even have fomo either.
    20. cedws · · focus · HN ↗
      Agreed, Sol has been my favourite since it released. I tried Opus 5 for a while and it made me want to throw my laptop out of the window.
    21. joduplessis · · focus · HN ↗
      Same. Sol was actually the reason I upgraded my plan to the $100 one. Hoping GPT6 Sol is the same.
    22. Imanari · · focus · HN ↗
      There are multiple models competing with 5.6sol on AA but none of them have the same feel (intuition,taste,judgement) - actually they are very far behind. I would say open source models are farther behind the the big labs than the benchmarks make you believe.
    23. ljm · · focus · HN ↗
      GPT does seem to stay out of the way and get things done. Only thing I notice is that the question tool&#x2F;elicitation doesn&#x27;t work that well any more so the thing doesn&#x27;t stop to wait for input.

      But I wonder if that&#x27;s intentional because it can keep computing while you are answering, so long as your steer aligns well enough with the direction it wants to go. Better than letting a cache go cold and burning compute on bringing it all back up.

  28. meerita · · focus · HN ↗
    OpenAI, Antrophic and others are operating with 80% margins. They can lower the prices for a long while.
    1. slekker · · focus · HN ↗
      Source?
      1. system2 · · focus · HN ↗
        Not a source but a comparison with a weaker non-SOTA model:

        Nvidia&#x27;s top AI chip Rubin sells in 72-GPU racks for about $3.5–7.8M. A rack running Xiaomi&#x27;s MiMo V2.6 Pro generates roughly 150–300B tokens a day, worth about $130–260k at Xiaomi&#x27;s API price. That&#x27;s a payback of the infrastructure in a few weeks in theory. After a few weeks or a month, the only cost is electricity, and whatever they make after that is pure profit.

        OpenAI and Anthropic are practically scamming people with the token prices.

    2. thimabi · · focus · HN ↗
      You probably mean they are operating with 80% margins discounting training expenses, which will continue to be pretty high for the foreseeable future.
      1. wyre · · focus · HN ↗
        Seeing how cheaply Xiaomi was able to train Mimo 2.6 I am starting to wonder if they are greatly over-exaggerating the real training costs to increase their valuations and investments.
        1. desterothx · · focus · HN ↗
          that was rl, not training. its a fraction of the cost
          1. anthonyrstevens · · focus · HN ↗
            Don&#x27;t let facts and knowledge get in the way of a good conspiracy theory!
  29. simianparrot · · focus · HN ↗
    Well at least it looks like OpenAI is dogfooding because their announcements, product names, and everything else looks and sounds like LLM-slop.
  30. sehw · · focus · HN ↗
    sage
  31. msh · · focus · HN ↗
    I dont understand why there is not a gpt-6 terra?
    1. blovescoffee · · focus · HN ↗
      It wasn&#x27;t really used enough and it sat in an awkward middle space between luna and sol where either luna high&#x2F;xhigh or sol med were better cost&#x2F;perf wise
      1. OutOfHere · · focus · HN ↗
        I don&#x27;t agree. In &quot;none&quot; thinking mode, Terra serves a useful purpose where high but not top-tier intelligence is needed. Luna doesn&#x27;t cut it.
    2. Tadpole9181 · · focus · HN ↗
      It probably didn&#x27;t see that much use, as it struggled to find a niche. If you wanted intelligence tasks, Sol was cheap enough and much smarter. If you wanted performance and cost-effectiveness, Luna was significantly better value while being only a little less intelligent.

      Terra ended up just being an awkward middle ground that was not particularly suited for any workload.

      1. OutOfHere · · focus · HN ↗
        That&#x27;s bs. The users of Terra disagree. Terra is useful when high but not top-tier intelligence is needed.
        1. Tadpole9181 · · focus · HN ↗
          Well, naturally people who use it are doing so because they like it. But I&#x27;m sure OpenAI is looking at the relative value&#x2F;usage of Terra next to the other tiers.
      2. msh · · focus · HN ↗
        I have found it worked quite well as the workhorse model in my hermes agent.
      3. darklinear · · focus · HN ↗
        I disagree. After a bit of experimenting, I actually found Terra to be a very good workhorse model on none-to-medium reasoning, and I actually quite prefer its code to Sol&#x27;s in many cases. It has less of a complexity to over-complicate things. Where Sol would have a sea of try&#x2F;except and recoveries for situations that are structurally impossible, Terra would just write nice, sequential code.

        Maybe for one-shotting large things Sol is better, but for prod code where I decompose into smaller tasks and read all the code I favored Terra.

        Sol 5.6 was still king for architecture&#x2F;research in my workflow, though.

    3. smith7018 · · focus · HN ↗
      I read that there are rumors that they&#x27;re getting rid of that tier. No idea where the rumor came from, though. This lends credence to it, I suppose.
    4. yawnxyz · · focus · HN ↗
      Terra was always worse of both worlds (expensive and not that good) so I think they&#x27;re just retiring it
    5. sandos · · focus · HN ↗
      I&#x27;m scared now, my employer only allows Luna and Terra on 5.6. I really hope they will allow Sol then on GPT 6.

      Funny thing is they very recently also set a real limit per-user&#x2F;month, so why even limit the models because theyre &quot;too expensive&quot;.

      1. apitman · · focus · HN ↗
        Your employer should reconsider. Sol high is cheaper than Terra max and smarter, when measured per task. ie even if tokens are more expensive Sol can often do a job with fewer tokens.
    6. miohtama · · focus · HN ↗
      Sol price is halved so no need for terra
  32. devinprater · · focus · HN ↗
    Good. Maybe they can use GPT-6 to fix the accessibility of their iOS app. Output shows as text fields to VoiceOver, and the accessibility announcements have backslashes before seemingly every punctuation mark. And then bring accessibility announcements to the Android app so I don&#x27;t have to make a whole new app just to add that through an accessibility service. Ugh the things I do for accessibility cause I&#x27;m blind. On a better note though, AI has done so much for the blind community, from image (and increasingly video) description to mods for video games like Final Fantasy 1 through 6 Pixel remaster, I have a ton to be grateful for.
  33. markerbrod · · focus · HN ↗
    Does anyone know if the ~50% price reduction also implies x2 subscription usage? Or is it only for the API.

    Edit: Yes, it applies also to subscriptions, source <a href="https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2102463847714247142" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2102463847714247142

    1. jrflo · · focus · HN ↗
      That’s great, I hate the opaqueness around subscription rates but at least it will show up in some way there too
  34. fHr · · focus · HN ↗
    Luna is the goat for real, cost intelligence ratio is insane already and it is enough for most daily computer use.
  35. yipinwong · · focus · HN ↗
    I&#x27;ve been raving about Luna 5.6 as it&#x27;s dirt cheap, and &quot;intelligent enough&quot;. Double quoted.

    Now GPT 6 Luna is even cheaper, and more intelligent, there is no going back... to SOL 5.6 for intelligent layer.

    1. dmazin · · focus · HN ↗
      Per the benchmarks in the post, Luna 6 is at best a couple points superior to Luna 5.6 and (unless I’m reading it wrong) xhigh has actually degraded in quality?

      I was hoping for a serious Luna upgrade. It was already cheap enough. This feels more like a price reduction than an upgrade.

      That said, if the new Luna is able to handle ultra mode and subagents v2 in codex cli, then at least that’s a win.

      1. yipinwong · · focus · HN ↗
        Benchmark doesn&#x27;t really show the whole story.

        I forgot which model degraded in quality as time went by, but let&#x27;s try out Luna 6 for a few more days to confirm for upgradability.

        1. tripledry · · focus · HN ↗
          &gt; Benchmark doesn&#x27;t really show the whole story.

          For me it seems like benchmarks are mostly noise, and the rest is based on vibes. Some find newer models annoying, some are amazed.

    2. elcritch · · focus · HN ↗
      If you put Luna on Max it&#x27;s still cheaper than Sol, but can achieve similar results. Though slower and with more iterations. Still it barely nudges my subscription usage!
      1. yipinwong · · focus · HN ↗
        ty for the suggestion. I really never used &quot;max&#x2F;ultra&quot; on Luna, and will give it a try.
    3. wartywhoa23 · · focus · HN ↗
      Ah, the ravers are not what they used to be anymore...
      1. cindyllm · · focus · HN ↗

        [dead]

      2. yipinwong · · focus · HN ↗
        Price is a big selling point for a normie like me.
  36. recitedropper · · focus · HN ↗
    This is the most blatantly astroturfed thread I have ever seen on Hacker News.

    My previous comment--which suggested astroturfing--was the highest upvoted comment here until it got flagged. Which implies to me that atleast the other remaining humans on this forum see it as well.

    26 minutes, 89 comments, upvoted instantly to the top, posted within one hour of the Opus 5.5 announcement. You tell me.

    1. john_strinlai · · focus · HN ↗
      &gt;My previous comment--which suggested astroturfing--was the highest upvoted comment here until it got flagged

      fyi, i flagged it because it is boring reading and against the rules.

      if you suspect astroturfing, flag the comments and contact the mods.

      (complaining that your complaint got flagged is also tiresome. contact the mods. &quot;@dang&quot; doesnt work, use the email.)

      1. recitedropper · · focus · HN ↗
        Do you think the &quot;you can&#x27;t mention astroturfing&quot; rule is really serving HN these days? Do you think this thread hasn&#x27;t been manipulated?

        I respect you for replying here though, and yes I get that HN forum standards would suggest flagging my previous comment. But it is just sad to see a place used to be so vibrant get manipulated because of how much weight it holds for us in the industry.

        And yea sure, I could go and flag all the bots and message Dang. But probably time to stop shouting into the void. :)

        1. seizethecheese · · focus · HN ↗
          [delayed]
          1. recitedropper · · focus · HN ↗
            More than two comments per minute right when it was posted, all simple one-liners celebrating the price decrease.

            I would agree now that the thread has recovered to a more interesting state, but how it first looked--combined with it being posted right after Opus 5.5 announcement--look questionable to me.

        2. minimaxir · · focus · HN ↗
          Yes, this thread is not being manipulated. People a) being excited about something and b) it being in favor of a certain company is not sufficient evidence of astroturfing.
          1. recitedropper · · focus · HN ↗
            I&#x27;ve seen your articles in the past and I thought they were great. So I respect your thinking, and have no intention of being combative.

            But do you really think people were so excited about cheaper versions of Astra that they were just waiting around to comment the instant this was posted? More than two comments per minute? All the initial comments were really similar too: brief one liners celebrating the cheap prices.

            I think AI right now is a sort of Rorschach test. What it is clearly revealing to me is that I don&#x27;t trust organizations with enormous financial incentives to not manipulate public opinion. So I see bots everywhere. :)

            1. minimaxir · · focus · HN ↗
              Yes. See my comment on Bluesky: <a href="https:&#x2F;&#x2F;bsky.app&#x2F;profile&#x2F;did:plc:oxaernim5mj2mmy3ytrvb42n&#x2F;post&#x2F;3mw4t3r2r5c25" rel="nofollow">https:&#x2F;&#x2F;bsky.app&#x2F;profile&#x2F;did:plc:oxaernim5mj2mmy3ytrvb42n&#x2F;po...

              &gt; what the fuck

              1. recitedropper · · focus · HN ↗
                I guess there are more people eagerly watching price decreases than I realized!
        3. john_strinlai · · focus · HN ↗
          &gt;Do you think the &quot;you can&#x27;t mention astroturfing&quot; rule is really serving HN these days?

          yes, i think so.

          because, unfortunately, complaining about bots (or astroturfing, or whatever) doesn&#x27;t stop them. so we end up with threads that have both the potential bot&#x2F;astroturfing&#x2F;whatever activity and complaints, which further drowns out any interesting comments.

          1. recitedropper · · focus · HN ↗
            That is actually a great point. I have no rebuttal.
        4. BenzeneDream · · focus · HN ↗
          Do you think all the comments in here are positive about the model? Because they aren&#x27;t. In no way does it seem astroturfed. And yeah, pretty boring to read that kind of comment every time.
          1. recitedropper · · focus · HN ↗
            Engagement is generally more important than purely positive sentiment.

            Anyway these comments were made when this thread was in an earlier state. I agree that it has gone on to be more &quot;organic&quot; looking. That doesn&#x27;t exclude it initially being manipulated to the top, in my mind, but certainly they aren&#x27;t carpet-bombing with only booster comments.

        5. FergusArgyll · · focus · HN ↗
          You have to accept that people are different than you.

          I see long massive pro apple threads. I don&#x27;t get it at all. As in; literally don&#x27;t understand what Apple is good for. But I have friends, family irl who love apple so I know the sentiment exists. I therefore accept that many HN users are similar.

          Many people really truly are happy to see another model drop and are excited about progress etc etc. Surely you&#x27;ve met such ppl in real life. Well, they&#x27;re here too (I&#x27;m one of them fwiw)

          1. recitedropper · · focus · HN ↗
            You know I think I have accepted this after a few decades on this planet, but probably can&#x27;t hurt to be reminded. :)

            I wouldn&#x27;t argue there aren&#x27;t real humans excited for this drop. It was just all the circumstances around it--the comment speed, the upvote speed, the initial uniformity of what people were saying.

            Anyway, thank you for the moral reminder.

    2. MaKey · · focus · HN ↗

      [dead]

    3. tomhow · · focus · HN ↗
      There&#x27;s no evidence of astroturfing. The comments you’re referring to are from accounts with established history and in different locations, without any evidence of being linked to OpenAI. They just seem excited about the models and the pricing.

      On the other hand, you have previously written: I&#x27;ll gladly admit I think what these companies are doing is unethical, and I&#x27;m sure that biases my thinking toward skepticism. [1]

      You have now posted accusations&#x2F;assumptions of astroturfing and manipulation at least 15 times, without ever providing any evidence. This is in breach of the guidelines, because comments like this poison discussions far more than the comments they&#x27;re complaining about.

      We – of course – want all comments and posts on HN to be authentic. HN is only a place where anyone wants to participate because since the beginning, we&#x27;ve had software mechanisms and moderation practices that detect and weed out inauthentic commenting and voting. We&#x27;re identifying and dealing with it every day, continually improving the software to detect and remove it. Most of that happens quietly and efficiently in the background without anyone having to see it. When users see evidence of manipulation and report it to us via email, we happily and thoroughly investigate it.

      Most of the time, what we find is simply that people are authentically excited and passionate about the topic, which is what is happening here. I understand it can be hard to accept that if you&#x27;re skeptical about the topic.

      It&#x27;s fine to be skeptical about the topic and you&#x27;re welcome to express your skeptical views on the topic. People do that every day on HN, about AI-related topics and countless others. Healthy debate is what we&#x27;re here for.

      But you can&#x27;t keep poisoning HN, by (1) continually posting these unfounded claims, then (2) when users and moderators simply uphold the guidelines, staging a protest by demanding your account be deleted. This is not what people do when they care about a forum&#x27;s health.

      [1] <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=48220908">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=48220908

      1. recitedropper · · focus · HN ↗
        Hopefully it is clear from my other comments that I do try to provide value too.

        &gt; You have now posted accusations&#x2F;assumptions of astroturfing and manipulation at least 15 times, without ever providing any evidence.

        I tried to point out the upvote speed and age-to-comment ratio for this thread look anomalous to me, and that it being posted within the Opus 5.5 release hour was further reason for skepticism. Circumstantial, sure, but I see very little ways to gather hard evidence of astroturfing without being a mod.

        &gt; when users and moderators simply uphold the guidelines, staging a protest by demanding your account be deleted.

        You&#x27;re right this was a little dramatic. I think it is just annoying though when two times that I have posted about astroturfing, it has been the most upvoted comment only to get flagged. I guess this is a self-fulfilling prophecy though, as you are right that other human commenters are abound and tend to flag people complaining about astroturfing.

        Anyway thanks for the reply. If your read of my account is that I&#x27;m more-often-than-not a bad actor, then I will stop commenting here. Seems like it is for the best. :)

        1. tomhow · · focus · HN ↗
          Like dang last week, it&#x27;s pleasantly surprising and welcome that you reply so cordially to our replies.

          I agree that you provide value, which is why we don&#x27;t just want to ban&#x2F;lose you.

          &gt; I have posted about astroturfing, it has been the most upvoted comment only to get flagged

          People love a conspiracy theory, and on a site like HN that has many people looking at it at once, it&#x27;s easy to get a large number of upvotes in a short amount of time if enough people find it exciting. We often see off-topic, titillating one-liners or ragebaity comments at the top of threads, and we always have to downweight them to keep the discussion on-topic and healthy.

          &gt; Anyway thanks for the reply. If your read of my account is that I&#x27;m more-often-than-not a bad actor, then I will stop commenting here. Seems like it is for the best. :)

          It seems like you&#x27;re well intentioned. You have your concerns about A.I., as many do and that&#x27;s fine. You&#x27;re still welcome here. Just please try to believe that many or most of the people who are enthusiastic about A.I. are as sincere in their positivity as you are in your concern.

    4. rvz · · focus · HN ↗
      Clearly this isn&#x27;t the first time. Remember criticizing model releases will get you flagged here.

      This is why HN has been on the down hill in quality and those that care to highlight that are being punished, while the astro-turfing, gaslighting and Show HN self-promotion slop continues.

      1. recitedropper · · focus · HN ↗
        A bunch of actual humans have replied to this so at least HN isn&#x27;t totally gone yet. :)

        Otherwise, yes, we agree. Although, given the other replies to this, there are clearly those who disagree who appear to be smart and level-headed.

        Anyhow I&#x27;ve learned my lesson now.

  37. jeffnash · · focus · HN ↗
    At this point, the deciding factors for me between Claude Code 20x and Codex Pro 20x are:

    1&#x2F; Usage limits: downstream of input&#x2F;output cost, but resets and obscure windows and odd 20x plan &#x2F; 5x plan != 4x usage math throw a wrench into it. Winner right now is Codex by a mile, especially when you factor in the fact that ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan. Always a bummer when asking if I should see a doctor about a rash means I can&#x27;t code as much. It&#x27;s also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems, giving better planning results or deeper code analysis without burning usage.

    2&#x2F; Context window in the harness. Claude Code wins on this. There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing (ETA: noname120 pointed out this is no longer the case and it can be enabled again [1]). 252k is just not enough. Codex&#x27;s compaction is very good, fwiw, but it happens so frequently that even a model as powerful as Astra sometimes loses the plot on long-running tasks.

    3&#x2F; Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.

    I&#x27;ve subscription hopped a bunch, and at times I&#x27;ve had both, but I keep coming back to Codex because it wins on 2&#x2F;3.

    ETA: apparently I haven&#x27;t been Keeping Up With the Altmans and new 20x signups have been disabled for a few weeks. I am grandfathered in, which makes the comparison above pretty much moot.

    [1]<a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49806060">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49806060

    1. spijdar · · focus · HN ↗
      I dunno about Codex-the-application itself, but you can definitely use e.g. Pi with the larger context windows with a Codex login. It puts a pretty large multiplier on credit usage, however.
      1. sidrag22 · · focus · HN ↗
        I&#x27;ve been doing this, my only experience with codex was brutal usage wise and i just retreated back to pi pretty quickly so the credit usage i&#x27;m receiving is kinda all im familiar with. Surely seems like less than CC, but i guess not using codex makes my experience kinda not valid for comparing usage.

        And ya i can go over that 240k limit, I still very seldom do, and try to treat it as the actual limit. I&#x27;m surprised to see so many people still talking about compaction to complete long running tasks, i think the bulk of the work should be somewhat frontloaded into a plan that is split off into subplans, then you can kinda open up a few options, one session with subagents for the subplans of the main plan, or just handoff prompts about progress against the main plan&#x2F;relevant subplan. I just never trust the blackbox that is compaction, I feel its a recipe for disaster&#x2F;context poison.

    2. noname120 · · focus · HN ↗
      &gt; It&#x27;s also is a godsend if you use an MCP like oracle to automate the process of calling the Pro model on particularly tough problems

      As far as I know Codex (at least the GUI) can automatically call the ChatGPT Chat models (including Astra 6 Pro), you just need to @ a ChatGPT Chat conversation from within Codex and tell it when to use it.

      &gt; There used to be a toml file workaround for Codex to extend the GPT context window to 1m, but this stopped working on the plans and only on per-token billing

      Not true, it works again[1]. I confirm that it works both on 5.6 Sol and Astra 6, possibly other models too.

      [1] <a href="https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2089082893804896524" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2089082893804896524

      1. jeffnash · · focus · HN ↗
        I actually haven&#x27;t played with the GUI. I probably should now that the Linux version is in beta. My situation is kind of the reverse: I like using oracle to basically zip up my repo, propose some sort of design based on the code, then provide a step by step implementation plan for a cheaper model to implement directly in a harness on my machine. I suppose I could do this and then save a step by referencing the oracle-created thread with the @ you mentioned

        And re: the toml workaround, AWESOME! I appreciate you pointing these two things out, this is my highest-ROI HN comment thus far.

    3. sodacanner · · focus · HN ↗
      In my personal experience I currently get a lot, lot more usage on the 5x Claude plan than the 5x Codex plan.

      Having limitless webUI ChatGPT usage is much better user experience, though. I&#x27;ll give them that.

      (edit: Sol-6 is half the price, so maybe the usage limits are going to be way better.)

      1. basisword · · focus · HN ↗
        I&#x27;ve been using Claude Pro and recently gave Codex a try again. Both on the $20 plans. I get so much more usage with Claude. It&#x27;s night and day for me. Codex runs out constantly, whereas Claude I hit limits very rarely.
        1. rbranson · · focus · HN ↗
          Assuming you are doing coding, I&#x27;m curious how would you characterize tne majority of your work (language, domain, frontend&#x2F;backend, etc)?
          1. basisword · · focus · HN ↗
            iOS development mostly. I&#x27;m using the Pro plans as it&#x27;s work on personal projects outside my day job and I&#x27;m able to get just enough usage from those plans to get me through each day.
        2. DaSHacka · · focus · HN ↗
          Same here, especially as I stick with Opus 4.6. My usage limits truly feel limitless, I can just hammer a task over and over again until completion.

          Meanwhile I just burned ~20% of my weekly quota with Astra making one config file for a service.

      2. jeffnash · · focus · HN ↗
        I&#x27;m actually interested to see how the token discount maps to the usage limit consumption. The conspiracy theorist in me wonders if they&#x27;re making up the discount and resultant load increase on the API end by reducing effective usage on the subscription end.
    4. paulmist · · focus · HN ↗
      &gt; Winner right now is Codex by a mile

      Opposite in my experience. I need to limit codex to 500k on medium&#x2F;low, still run out in 2-3 days with 1 CLI window. CC gives me 4-5 medium&#x2F;high days with 2-3 CLI windows, and Opus is still great for other regular dumb engineering&#x2F;refactoring.

      On the other hand my head starts to hurt if I read Opus for too long, hopefully they fixed it with 5.5.

      1. joshstrange · · focus · HN ↗
        This is my experience. After months of hearing how Codex limits were way higher I bumped to the $100&#x2F;mo plan after hitting my limits a day early on Claude due to some heavy usage + Fable (not normal for me, I often fit nicely in the $200&#x2F;mo plan). I hit the usage limit in a day with a single agent running on codex and the tiny context window was stifling. Yes, I&#x27;m comparing a $100 to a $200 plan but I extrapolated the usage (4x&#x27;d it) and it still wasn&#x27;t close, I got way more done with Opus.

        Using Agentsview (which might have it&#x27;s own issues) I was getting ~$200 of API usage in my 1 week Codex window (paid $100) vs ~$5,000 of API usage in 1 week for Claude (paid $200).

    5. lifty · · focus · HN ↗
      What are you talking about? ChatGPT unmetered? No way! That was 2 months ago perhaps and it’s possible your account still hasn’t gotten the new limits. I noticed around 1 month ago I was still going full throttle on my codex subscription and my limits were barely budging, and then all of sudden people around me started to complain about limits. I thought they’re crazy, but then my account go the hammer, and that was it. If I have the same pattern of usage like I did before, basically having an agent working continuously on a coding take, my weekly limit goes in 2 days.
      1. nwienert · · focus · HN ↗
        [delayed]
        1. lifty · · focus · HN ↗
          It’s all anecdotal based on my experience and other countless discussions I have seen online. I’ve heard speculation that once they hit 20 million codex users capacity is tighter so they have to manage it. The previous limits were unsustainable compared to token pricing.
      2. jeffnash · · focus · HN ↗
        On usage in ChatGPT settings, I see: Plan limits Shared across Codex, Work, Workspace Agents, and ChatGPT for Excel. Chat conversations are not included.

        Is this not the default anymore? I am on the (now closed) 20x plan.

        1. lifty · · focus · HN ↗
          That’s the default. I didn’t express myself clearly but I thinking your situation is not the common case anymore, or perhaps you are not using it hard enough. Codex limits deplete very fast these days, it’s not “unlimited”.
          1. jeffnash · · focus · HN ↗
            I am saying that because ChatGPT usage is unlimited, I don&#x27;t have to eat into my Codex limits when I use ChatGPT. Codex certainly has limits. Last time I had a Claude sub (hedging here since much of the info in my comment was outdated), my usage limits on claude.ai threads was shared with Claude Code.
            1. lifty · · focus · HN ↗
              Finally got the nuance. Indeed the chat part of the subscription is unlimited as far as I know. Now that part of your comment makes sense!
              1. jeffnash · · focus · HN ↗
                Sorry about that, I accidentally a word (hope that reference doesn&#x27;t date me)
    6. NorwegianDude · · focus · HN ↗
      Codex&#x2F;ChatGPT Pro 20x isn&#x27;t really a thing now, they have disabled it a week or two ago.
      1. yzydserd · · focus · HN ↗
        It’s back available since 4 days ago.
        1. malshe · · focus · HN ↗
          It&#x27;s not as of 4 minutes ago: <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;pricing&#x2F;" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;pricing&#x2F;
          1. jwitthuhn · · focus · HN ↗
            Not advertised, but for me it shows up in the app if I look at the available plans from my $20&#x2F;mo account.
    7. kornelijus · · focus · HN ↗
      Not that I disagree that Codex wins out, but the deciding factor actually is - Codex Pro 20x is not available for purchase, indefinitely. So, what&#x27;s the point of this discussion? People who already have the 20x sub are unlikely to cancel, and the rest of us can&#x27;t access it.
      1. lawgimenez · · focus · HN ↗
        I purchased mine using Apple&#x27;s in-app subscription. I just checked and it is still there.
    8. marcd35 · · focus · HN ↗
      theres a popular thread on claudecode or claudai subreddit that proves 20x isnt really 20x. apparently its a marketing gimmick and the recommended solution is two 5x plans &gt; 20x at greater than half the cost of the 20x
    9. glub · · focus · HN ↗
      &gt; Usage limits [...] Winner right now is Codex by a mile

      This hasn&#x27;t been the case since around July. If you measure usage in raw api costs, Anthropic is actually giving more on $200 than OpenAI now. This includes resets. Usage allocation difference would be humiliating for codex subs were it not for resets. But fixing usage limits with resets is ugly, and they&#x27;re not good for your mental well-being.

      &gt; Context window in the harness

      Codex now allows 1M for subs with config params. But generally speaking, you shouldn&#x27;t really be using 1M context. If you accidentally send a request with say, ~700k context already accumulated in a session which is outside cache TTL, you&#x27;re paying full cost of these 700k tokens.

      &gt; I&#x27;ve subscription hopped a bunch

      OpenAI actually has a new strategy to prevent subscription hopping after their 2-3 month-long marketing push to get claude-folks to switch over:

      you can&#x27;t buy a $200 sub anymore. So if you cancel, you won&#x27;t be able to get back in. Hostage situation, essentially.

      EDIT: re: usage limits, oh-my-pi maintainer has been tracking this - <a href="https:&#x2F;&#x2F;nitter.xitter.cc&#x2F;_can1357&#x2F;status&#x2F;2090075496948060372" rel="nofollow">https:&#x2F;&#x2F;nitter.xitter.cc&#x2F;_can1357&#x2F;status&#x2F;2090075496948060372

      1. ipsod · · focus · HN ↗
        &gt; you can&#x27;t buy a $200 sub anymore

        Are you sure?

        1. glub · · focus · HN ↗
          Yes.

          <a href="https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2098113585683808624" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;thsottiaux&#x2F;status&#x2F;2098113585683808624

          1. spiderice · · focus · HN ↗
            That is an old tweet. They since reenabled it. I know because I was on the $200&#x2F;month plan and couldn&#x27;t resub once it expired. However, a couple days ago it finally let me resub again.

            Now, if they disabled it yet again, that&#x27;s another story. But that tweet is not evidence of that.

            1. cthalupa · · focus · HN ↗
              I have been attempting to get on the $200 sub for a while. It was not available for me a few days ago, and checking again now, it is still not available.
              1. spiderice · · focus · HN ↗
                That&#x27;s too bad. I wonder why I was able to get it after days of not being able to. They must&#x27;ve just temporarily enabled it again. Probably worth checking a few times a day to see if it reappears.

                Though with the price of GPT-6 Luna, the temptation to switch to pay-per-token grows.

                1. coderenegade · · focus · HN ↗
                  You can resub on that plan if you&#x27;ve been on it before. They aren&#x27;t taking new subs on that plan for the time being.
                2. phil21 · · focus · HN ↗
                  There was&#x2F;is a loophole where if you signed up via the iOS or Android app, it allowed it.

                  It&#x27;s been disabled for some time now though otherwise, I check about once a day myself and keep and eye out on social media.

                  Annoying since I was about to upgrade back to the $200 plan after downgrading to the $100 plan due to being on leave and not needing as much usage the month prior. Doh.

            2. glub · · focus · HN ↗
              I think what they did was allow resubs for users who already had $200 sub before.

              Just checked my toy chatgpt account that only ever had a $20 sub. $200 plan still shows &quot;The 20X plan is temporarily unavailable for purchase&quot;.

      2. rudedogg · · focus · HN ↗
        I’ve been a Claude user, switched to Codex expecting usage limits to be more loose but I can’t even get through a basic sysadmin task on the $20 plan using Sol medium before I hit the 5hr one.

        I think I’m gonna move back to a Claude plan. I could barely hit the $200 limit if I went non-stop on programming tasks.

        1. qlte · · focus · HN ↗
          I do the bulk of work on Sol Medium&#x2F;Low and don&#x27;t have that experience on the $20 plan. If you said Astra I&#x27;d agree it&#x27;s easy to burn through the 5 hours even on the lower reasoning levels.

          Do you have &#x2F;fast enabled by any chance?

          1. rudedogg · · focus · HN ↗
            I don’t think so, I’ve seen it suggest I try it.

            Yes, I was considering the $100 plan, but I hit the 5hr limit in an hour, thought about how even at $100 I cant go non-stop on a single agent running Sol Medium and decided I should probably get back on Claude

            1. hirvi74 · · focus · HN ↗
              Sorry if I am misunderstanding you, but I am pretty sure the $100 plan doesn’t have a 5hr usage limit. So, if that was what was preventing you from going non-stop, it might be worth it.

              I am considering the plan myself. I just don’t know if I want to fork out $100 per month for something I will make $0 off of.

              1. rudedogg · · focus · HN ↗
                Thanks for the info, I didn’t know that about the 5hr limit
        2. glub · · focus · HN ↗
          I think OpenAI essentially executed a bait-and-switch here, and they&#x27;ve lost a lot of goodwill with me, like Anthropic did, before them.

          When they started the aggressive campaign, entire X (including myself, sadly) was full of posts about how &quot;unlimited&quot; codex usage is even on a $20 plan. Sam Altman was posting something in line of &quot;we love our users, unlike Anthropic&quot;. Got my network to get codex subs because of the value compared to claude.

          Then they gradually reduced the limits to the point where even $200 plan only lasts you just 1-2 days and $20 is basically unusable, then the hostage thing.

          1. ivm · · focus · HN ↗
            &gt; even $200 plan only lasts you just 1-2 days

            I’m working on 2–3 apps at once and barely manage to use 70–80% of the weekly quota, with everything being done on Sol Xhigh and, lately, all the planning on Astra. Still have a reset stashed too.

          2. malshe · · focus · HN ↗
            I remember Tibo Sottiaux telling people on X how OI doesn&#x27;t believe in 5 hour limit just a day or two before OI adopted it.
            1. phyrex · · focus · HN ↗
              tbf that&#x27;s only for the pro plan, not the two max plans
          3. Numerlor · · focus · HN ↗
            I think the models getting dumber impacted that too, after a couple weeks both sol and Luna felt notably worse to me than they did at release
          4. slopinthebag · · focus · HN ↗
            &gt; Then they gradually reduced the limits to the point where even $200 plan only lasts you just 1-2 days

            what? im on the $100 plan and ive literally never run out of usage, and thats mostly running Astra high.

            maybe its the harness

        3. hadlock · · focus · HN ↗
          I&#x27;ve run into hitting limits on the personal plan perhaps twice since the beginning of the year. But also I don&#x27;t use the personal plan for coding tasks between 7am-noon M-F.
        4. boardwaalk · · focus · HN ↗
          similar here: I tried Codex $20&#x2F;mo on a trial and I ran out of 5hr usage mid way through a medium complexity task on a medium size model twice and gave up there. I don’t recall the equiv Claude plan being anything like that. Anecdata, but not great for OAI if they actually want to retain people on a trial.
          1. cromka · · focus · HN ↗
            You don&#x27;t get Fable on Claude 20 USD plan. You get Sol on equivalent Codex plan.
            1. istjohn · · focus · HN ↗
              You meant Astra, not Sol, I think. But Opus 5.5 is slightly better than Fable and Astra now.
              1. cromka · · focus · HN ↗
                Yes, it&#x27;s Astra and Sol before.

                Opus 5.5 is better in benchmarks, but has substantially less parameters so is world knowledge cannot compare against Fable or Astra.

        5. this_user · · focus · HN ↗
          Astra is barely usable even on the $100 plan. And that is if it doesn&#x27;t just burn through 80% of your weekly quota in a couple of hours by continually expanding the scope of the task you gave it - while not noticing the failing tests that are right in front of it.

          Opus is at least actually usable even on the small plan. The main downside is its insane writing style, but 5.5 seems to address that somewhat. Otherwise, you can just use your $20 OpenAI plan to have Luna de-slop Opus&#x27; prose, which seems to work fine.

          1. Huppie · · focus · HN ↗
            I have a Claude Code hook that calls codex for a code review on commit time (Codex is set to Astra Medium) and it&#x27;s been pretty good in general. It sometimes hits the 5hr limit but most of the time it provides really good feedback and because it&#x27;s a completely different model it&#x27;s mostly complementary to what Fable&#x2F;Opus do themselves. IMHO it&#x27;s been $20 well spent.

            ...but the few times I&#x27;ve tried to use codex for a moderately difficult task it burned through its limit extremely quickly.

        6. cromka · · focus · HN ↗
          But you don&#x27;t get Fable on Claude 20 USD plan, then why compare it Sol on Codex 20 USD?
          1. rudedogg · · focus · HN ↗
            Sol is their middle model. Luna is smallest. And Astra is big, their Fable equivalent.
            1. matheusmoreira · · focus · HN ↗
              My code review benchmark put Sol 5.6 on the same performance tier as Fable 5.

              <a href="https:&#x2F;&#x2F;www.matheusmoreira.com&#x2F;articles&#x2F;code-reviewing-lone-lisp-with-sol-and-fable" rel="nofollow">https:&#x2F;&#x2F;www.matheusmoreira.com&#x2F;articles&#x2F;code-reviewing-lone-...

          2. sisyphus15 · · focus · HN ↗
            Sol is OpenAI&#x27;s Opus, and Astra is OpenAI&#x27;s Fable. Both pricing-wise, and performance-wise.
        7. joquarky · · focus · HN ↗
          On the $20 plan, you can&#x27;t use Sol for much more than planning and review. Luna xhigh for the rest. Have Sol write the plan specifically for Luna so it adds more direction and validation to the plan.
      3. platinumrad · · focus · HN ↗
        Given that Anthropic models are very verbose and OpenAI models can be very concise, wouldn&#x27;t a count of expected task completions be a better measurement than raw API costs?
        1. glub · · focus · HN ↗
          Perhaps. But Sol&#x2F;Astra also likes dumping pages of jargon-packed content at me, so I&#x27;m not sure it&#x27;s that much different. I actually still prefer the way Fable talks to me, even considering the horrible claudisms.

          But even if we leave that aside, OpenAI models are also much more eager than Anthropic, which are on the lazier side. Left unsupervised, Sol&#x2F;Astra will attempt to build a sha256 verified rocket ship if you ask them to fix a race condition in your to-do list app. Anthropic models will do what you asked for, maybe even forget to implement parts of that ask, but they won&#x27;t generally throw a slop granade at you.

          I can leave Fable orchestrator unsupervised for ~2h. Leaving Sol&#x2F;Astra unsupervised for ~2h means the next user turn will contain a message: &quot;what are you doing and why?&quot;.

      4. InsideOutSanta · · focus · HN ↗
        I think the problem with Anthropic&#x27;s plan is that Fable just destroys it. If you stick to Opus and below, the $200 plan goes from &quot;using 50% of the weekly quota on the first day&quot; to something much more reasonable.
      5. jrflo · · focus · HN ↗
        Do you have a source on the first note? I switched away from Claude around July because of how bad the usage limits were, and Codex gave me easily double the amount of usage per task completed. Would be interested to see if that&#x27;s no longer the case.
        1. glub · · focus · HN ↗
          Added link in edit. OMP maintainer has several claude and codex subs and he&#x27;s been tracking usage since around July.

          I haven&#x27;t been tracking, but this roughly matches my experience with codex 20x and claude 20x subs. Claude subscription now lasts me 3-3.5 days on average. Codex is 2-2.5 days. This is work on same projects, with similarly sized tasks.

          To make matters worse, I&#x27;ve merged a lot more code produced by fable than sol&#x2F;astra.

      6. athrowaway3z · · focus · HN ↗
        I&#x27;m not sure the tokens can be compared like that between OpenAI&#x2F;Anthropic.

        When i swapped between a 200k Fable context into an Astra model (i was out of fable) the token usage in that context dropped to 150k or something.

        Either there was a bug somewhere, or the same text got cut up very differently between providers.

        1. glub · · focus · HN ↗
          That 50k was almost certainly accumulated encrypted reasoning tokens that would have been unreadable by astra.
          1. athrowaway3z · · focus · HN ↗
            Ah that makes sense.
      7. cameronh90 · · focus · HN ↗
        To add my anecdote, while the Codex subscription appears to get you much fewer tokens as measured by cost, I find the amount of actual useful work that can be done by both subs to be about equal. Codex seems much less prone to burning millions of tokens just reading the codebase and doing nothing useful. That also makes it much quicker. Plus it actually does what I tell it with few mistakes first time, so less rework needed.

        The Claude TUI is just so much better though so I&#x27;m hoping Opus 5.5 is actually good and not just benchmaxxed.

      8. matheusmoreira · · focus · HN ↗
        Anthropic has a separate meter for Fable. I used to get like five Fable sessions per week and that&#x27;s it.

        OpenAI has no such nonsense. No separate meter. No five hour limits. I get to use Astra at max effort on literally every task if I want to, and even this somehow lasts me several days.

        Anthropic got caught playing stupid &quot;20x refers to the 5h limit&quot; word games with their customers. Meanwhile, I have statistically verified that OpenAI Pro 20x = 4 * Pro 5x = 20 * Plus, exactly as advertised.

        I quantified cybersecurity lockouts on my code review benchmark and they were significantly lower on OpenAI:

        <a href="https:&#x2F;&#x2F;www.matheusmoreira.com&#x2F;articles&#x2F;code-reviewing-lone-lisp-with-sol-and-fable" rel="nofollow">https:&#x2F;&#x2F;www.matheusmoreira.com&#x2F;articles&#x2F;code-reviewing-lone-...

        My benchmark also suggests even OpenAI&#x27;s Sol models can match Fable performance at a fraction of the cost.

        OpenAI also used to have a ton of very nice features: unlimited chat separate from codex, allowing turns to finish even at 0% usage remaining. Sadly these got removed after abuse.

        As a former Anthropic customer, OpenAI is simply the better company. There is no way around it. Good place to be while the chinese open weights models catch up. Claude is good but it doesn&#x27;t make up for Anthropic&#x27;s shenanigans.

        1. ghostpepper · · focus · HN ↗
          OpenAI has 5 hour limits on the $20 plan. I agree about cybersecurity refusals though.
      9. albert_e · · focus · HN ↗
        &gt; If you accidentally send a request with say, ~700k context already accumulated in a session which is outside cache TTL, you&#x27;re paying full cost of these 700k tokens.

        Thinking aloud:

        The harness UI should probably implement a timer that shows whether you are still within Cache TTL since your last turn of the conversation.

    10. joshstrange · · focus · HN ↗
      Maybe it&#x27;s due to 20x &#x2F; 5x != 4 but I have the $200&#x2F;mo Claude and $100&#x2F;mo Codex and I get _way_ less usage on Codex, well under 1&#x2F;4th the usage. In 1-2 days of semi-heavy _single_ agent usage with Sol High I can burn through my whole week of Codex. Again, this is not running multiple agents, just 1 at a time.

      Compare that to Claude and I can run multiple agents on Opus almost indefinitely. YMMV of course but I was shocked at how quickly I burned through Codex usage.

      On the context window, I feel so cramped on Codex, compacting happening every time I turn around is annoying. I didn&#x27;t realize how much I enjoyed the Claude context window size.

      1. rgbrenner · · focus · HN ↗
        Same experience. Have both subs. It&#x27;s just not true anymore that Codex gives you more usage than Claude.

        Makes me think they picked Codex, stopped trying Claude, and just hang on to outdated beliefs about the value they&#x27;re receiving.

      2. hirvi74 · · focus · HN ↗
        Isn’t that to be expected when comparing one 20x plan to another 5x plan?

        I am curious how the 5x plans differ between both providers.

        1. malshe · · focus · HN ↗
          I have 5x on both of them. I get way more use from CC than Codex. Actually as we speak, I exhausted my Codex limit twice in the last two days. I am living on banked resets right now.
    11. rgbrenner · · focus · HN ↗
      &gt; Claude Code 20x and Codex Pro 20x

      That isn&#x27;t a valid comparison, since Codex 20x is closed. So we should be comparing Claud 20x to Codex 5x + credits.

      1. wahnfrieden · · focus · HN ↗
        You’re sharing outdated info
        1. rgbrenner · · focus · HN ↗
          Care to be more specific? 20x is closed.
        2. rgbrenner · · focus · HN ↗
          Care to be specific? 20x is closed. And the 2x pricing is literally on the pricing sheet for gpt-6 astra, sol and luna.
          1. wahnfrieden · · focus · HN ↗
            The info you’re citing is for API not Codex
    12. [deleted] · · focus · HN ↗

      [deleted]

    13. elxr · · focus · HN ↗
      Also, OpenAI is just a company I&#x27;d rather support than Anthropic.

      While you&#x27;re understandably not including the values of the $20 standard plans on both, I find the generosity of then token limits on ChatGPT plus vs Claude Pro (it&#x27;s a huge difference) to be good representation of their respective attitudes towards the average user. You literally cannot use Claude pro to build real software, unless you&#x27;re extremely frugal with your prompts and don&#x27;t try anything even a little ambitious.

      Also, Anthropic has zero models comparable to Luna.

      1. bix6 · · focus · HN ↗
        Reasons for this?

        &gt; Also, OpenAI is just a company I&#x27;d rather support than Anthropic.

        1. elxr · · focus · HN ↗
          Their responses towards using their subscriptions on opencode for one. Second, Dario just has a habit of making completely doomer comments on the future of software engieering as a job and towards the open-weights model ecosystem.

          Sure, he&#x27;s free to say whatever especially considering the amount of revenue he&#x27;s creating, but it&#x27;s just an altitude that I prefer not to see.

          1. bix6 · · focus · HN ↗
            And Sam is better?
            1. elxr · · focus · HN ↗
              Significantly.
            2. CuriouslyC · · focus · HN ↗
              Sam is sketchier on a personal level, but judged just on the words coming out of their mouths, he&#x27;s also much less paternalistic&#x2F;controlling and more customer focused.
              1. felixgallo · · focus · HN ↗
                I think any amount of &#x27;paternalistic&#x2F;controlling&#x27; turns out to have been justified when, after dismantling the safety teams and pretending not to know what safety is, OpenAI had the HuggingFace series of scandals. You can dislike the idea of safety and people talking about safety, but not only is the evidence right there, but OpenAI came out shamefacedly and literally agreed with Amodei&#x27;s statements, including that they agreed to pace the frontier.
                1. platinumrad · · focus · HN ↗
                  Your entire recent comeback history is defenses of Anthropic. Do you work for them?
                  1. felixgallo · · focus · HN ↗
                    It isn&#x27;t, and no.
          2. hbrn · · focus · HN ↗
            I think opencode subscription issue is just a different marketing strategy. Neither company wants it, but OpenAI believes it&#x27;s worth it as a marketing expense in the long run.

            And Dario&#x27;s &quot;AI will kill us all&quot; is the same as Sam&#x27;s &quot;AI will discover ALL science and we&#x27;ll be building Dyson spheres&quot;.

            Different flavors of the same BS.

            1. platinumrad · · focus · HN ↗
              The first one terrifies people who really don&#x27;t need to be. It&#x27;s deeply unethical.
          3. usef- · · focus · HN ↗
            I think if they truly believe it&#x27;s happening we generally want to encourage them to be honest with the public, though, don&#x27;t we? We&#x27;ve spent decades complaining about ceos not being honest in the public risks that they see
            1. killingtime74 · · focus · HN ↗
              Last week he said they should pause research and today there just released newer and better models. His talk is completely meaningless.
              1. usef- · · focus · HN ↗
                He didn&#x27;t say they were pausing research. You might have only read the social media responses to his essay, not the essay itself. Social media seems even less accurate than usual when it comes to anything AI related.
      2. therein · · focus · HN ↗
        They are both companies I&#x27;d rather not support. Not that our support for them has any material impact. NVIDIA is bankrolling them directly and indirectly.
      3. InsideOutSanta · · focus · HN ↗
        &gt; Also, OpenAI is just a company I&#x27;d rather support than Anthropic.

        They&#x27;re both pretty horrible, but I find it difficult to find arguments for why Anthropic is worse than OpenAI, other than their doomtrolling. Which, in the grand scheme of things, doesn&#x27;t even register.

        Edit: forgot about the SpaceX thing.

        1. elxr · · focus · HN ↗
          OpenAI has been way more open with users using their subscription plans on 3rd party tools.

          That alone is reason enough. Also, I don&#x27;t think either of them are horrible. That&#x27;s honestly a ridiculous take considering how much people in here love their models, and how much they&#x27;ve advanced the industry forward.

          1. InsideOutSanta · · focus · HN ↗
            &gt; Also, I don&#x27;t think either of them are horrible. That&#x27;s honestly a ridiculous take considering how much people in here love their models

            That&#x27;s a non-sequitur.

            &quot;Nestle is a great company, considering how much people love their chocolate.&quot;

            1. elxr · · focus · HN ↗
              How about you tell me what makes the horrible then. There&#x27;s pluses and minuses to both obviously, almost everyone around me have positive experiences with the product. They&#x27;ve innovated at a pace unheard of before 2026, and for openAI specifically the amount of value they&#x27;ve provided to me and family members (who aren&#x27;t even developers in the slightest) has far outweighed the supposed horrible actions they&#x27;ve done.

              Yeah I don&#x27;t think the handling of copyrighted training data was correct, but I can&#x27;t pretend I know what the correct solution to that issue is.

              Speaking of OpenAI specifically, they don&#x27;t price gouge people, they aren&#x27;t aggressively anti-competitive, they&#x27;re not nearly the perpetual hypocrisy machine that Anthropic is (which is one thing I actually really dislike).

              Regarding Nestle, it&#x27;s pretty obvious that the sentiment towards them is a lot more negative and they aren&#x27;t universally loved by any group of people. Processed foods are by and large garbage nobody needs. Their use of forced labor is denounced by just about everyone. What have OpenAI&#x2F;Anthropic done that&#x27;s even similar in scope to the forced labor &#x2F; modern slavery that people hate Nestle for.

              If you had a company that genuinely helped hundreds of millions of people worldwide become more productive and more satisfied with their tools, and the overall sentiment towards your products within the industry is positive, then what argument would there be that your company is &quot;horrible&quot;? At least give some decent counter arguments.

              1. mullingitover · · focus · HN ↗
                &gt; What have OpenAI&#x2F;Anthropic done that&#x27;s even similar in scope

                You mean aside from &quot;the largest theft of labor in human history&quot;[1]?

                [1] <a href="https:&#x2F;&#x2F;www.nytimes.com&#x2F;2026&#x2F;09&#x2F;17&#x2F;technology&#x2F;microsoft-openai-publishing-industry.html" rel="nofollow">https:&#x2F;&#x2F;www.nytimes.com&#x2F;2026&#x2F;09&#x2F;17&#x2F;technology&#x2F;microsoft-open...

          2. jsw97 · · focus · HN ↗
            For me the first point, openness to 3rd party, is the decider. I don’t want to build tooling around a completely closed model. I liked being able to use pi, and now I exclusively use my own harness which I modify the way I want. Not possible with Anthropic subscription.
            1. elxr · · focus · HN ↗
              100% agree.

              I often have the urge to design my own harness too (once I have more time). But even with the current mainstream harnesses out there, there&#x27;s just to many hurdles if you wanted to mainly stick with anthropic models and need the subsidized pricing (from a sub).

          3. ychnd · · focus · HN ↗
            They are both killing people &#x2F; aiming for murder bots, aren&#x27;t they?
            1. ketzu · · focus · HN ↗
              &gt; aiming for murder bots

              Anthropic prohibits the use of claude models for development of lethal technology afaik eg [1].

              [1] <a href="https:&#x2F;&#x2F;www.epc.eu&#x2F;publication&#x2F;the-pentagon-blacklisted-anthropic-for-opposing-killer-robots-europe-must-respond&#x2F;#:~:text=Anthropic%20a,and%20mass%20surveillance." rel="nofollow">https:&#x2F;&#x2F;www.epc.eu&#x2F;publication&#x2F;the-pentagon-blacklisted-anth...

              1. ychnd · · focus · HN ↗
                But only in a fully autonomous way, they have nothing against mass surveillance outside of US and killing non-autonomously... I feel like that&#x27;s also quite a low bar, but apparently enough for them to get bullied.
        2. andriy_koval · · focus · HN ↗
          &gt; why Anthropic is worse than OpenAI

          Anthropic is trying to kill open models way harder

        3. nullc · · focus · HN ↗
          OpenAI just wants to make money, perhaps through underhanded tactics if they can get away with it.

          Anthropic does all that but they&#x27;re also populated by many people who believe they are building God and that they must build their god first in their own image so that it can take control of humanity and protect us from any competing god which is not built in their image. Their position is inherently paternalistic and authoritarian, and they consider suppression of competition not just important to the bottom line but to life in the universe. Under the doomer ethos there is no evil too great to rationalize.

          There are plenty of wrongs done in the name of profit, but capitalists have nothing on zealots in terms of causing serious harm. Profit motives can be directed by influencing incentives, but zealotry is frequently terminal.

          That isn&#x27;t to say that there isn&#x27;t some overlap-- the cultists have infected both organizations. But OpenAI has pretty consistently only given lip service to AI doom to the extent that it improves the bottom line, while (mis)Anthropic was founded specifically because OpenAI wasn&#x27;t mentally ill enough.

          1. elxr · · focus · HN ↗
            Well said. The superiority complexes from the Anthropic messaging on their presentations&#x2F;blogs&#x2F;articles is just too much, even for a frontier AI company.

            Anthropic has great products, but it&#x27;s not meaningfully better to 99% of devs that I&#x27;d rather support the company that doesn&#x27;t constantly act in opposition to optimism and to the vibe I&#x27;d prefer for a 100 billion dollar (or however ridiculous amount they&#x27;re worth now) tech company embraces.

            AI doomerism is a genuine waste of time if you aren&#x27;t actively pushing towards a better AI industry for everyone, not just the groups in full ideological alignment to your personal leanings.

      4. felixgallo · · focus · HN ↗
        You&#x27;d rather literally support &lt;i&gt;Sam Altman&gt;&#x2F;i&gt;? I mean, that&#x27;s a position to take, for sure, but apparently several people still use Grok, so maybe it&#x27;s not all that surprising.

        &quot;You literally cannot use Claude pro to build real software, unless you&#x27;re extremely frugal with your prompts and don&#x27;t try anything even a little ambitious&quot; - that&#x27;s way past ridiculous. Even just using Fable most of the time, working on several ambitious projects, I have a hard time hitting the limit with a Max plan.

        1. qwerpy · · focus · HN ↗
          Lol Grok users catching strays here. I enjoy it and it has built some nice things for me as a hobbyist. The attitude of the company is more just quietly build cool things rather than Anthropic&#x27;s holier-than-thou condescending attitude coupled with the over the top self-serving doomerism.
      5. ketzu · · focus · HN ↗
        &gt; You literally cannot use Claude pro to build real software

        Interestingly I would have drawn the exact opposite conclusion looking at my Claude and codex usage.

        I can&#x27;t get anything sustained out of codex in chatgpt plus, while I have been using Claude pro extensively and put on a lot of experimental task and features.

        I ran into codex exhausting a 5h window on code review in minutes (like 3minutes) multiple times, while I could get Claude to implement 2~3 medium sized features with the same usage consumption.

        (I also really dislike the usage resets in codex, they always make me feel like I use them wrong because I often just want to reset the 5h window, but they can only do both at once...)

    14. rbranson · · focus · HN ↗
      Astra planner&#x2F;designer with Sol+Luna subagents has worked well for me to improve context continuity. Luna generates code, Sol reviews code and runs&#x2F;monitors integration&#x2F;E2E tests. It&#x27;s about 20% more usage efficient and 20% faster to finish tasks. I&#x27;ve been very subagent-skeptic for a while but the economics of codegen with Luna have made it click. This just works in Codex with a single-line AGENTS.md instruction.
      1. NolF · · focus · HN ↗
        Do you mind sharing? I would love to give it a try and see if I can stretch the x5 plan further.
      2. pyinstallwoes · · focus · HN ↗
        How do you do this?
        1. vatsachak · · focus · HN ↗
          Literally tell the model, spawn a &lt;model_n&gt; to do &lt;task_n&gt; and it will do it
    15. hintymad · · focus · HN ↗
      &gt; Ability to use the plan outside of the official harness. Codex wins. Anthropic does shit like bills requests as extra usage if it sees a hermes.md in a commit.

      I&#x27;m quite puzzled about why Anthropic is so hellbent on blocking other coding agents. It&#x27;s not like Claude Code has any secret sauce, right? And doesn&#x27;t Anthropic make monkey off API usage, and their magic is on the model side anyway?

      1. InsideOutSanta · · focus · HN ↗
        They want to lock people into using the Claude Code ecosystem to make switching to other providers more difficult.
      2. glub · · focus · HN ↗
        It&#x27;s for lock-in - same reason why it took them so long to finally support AGENTS.md.

        But to be fair, they don&#x27;t really enforce the harness rule that much anymore. I guess if your harness doesn&#x27;t do a lot of weird things like a lot of cache misses, or triggers some distillation attacks, or some broader Chinese fingerprints, they&#x27;re tongue-in-cheek okay with you using a third party harness.

        1. codybontecou · · focus · HN ↗
          You can use Claude’s subscription in Pi now? Last I tried it opted for extra usage.
          1. glub · · focus · HN ↗
            Not natively, as it&#x27;s still a ToS violation and adding that in pi would go against pi principles, but there are many plugins&#x2F;proxies that make it work.

            oh-my-pi supports it natively (again, still a ToS violation), by impersonating claude code&#x27;s fingerprints.

            I have been using oh-my-pi with 3 claude subs for the past few months without any issues. Even native server-side OAI&#x2F;ANT compaction works out of the box.

            1. goosejuice · · focus · HN ↗
              [delayed]
          2. manx · · focus · HN ↗
            Yes, you need an extension that uses the claude code credentials from the file system, like this: <a href="https:&#x2F;&#x2F;github.com&#x2F;fdietze&#x2F;pi-claude-auth" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;fdietze&#x2F;pi-claude-auth

            Works pretty well for me, even with latest opus-5-5

      3. nl · · focus · HN ↗
        Originally it was because Anthropic was so compute constrained they relied on the extra care the Claude harness took with caching (heavy use of cache breakpoints etc) that other harnesses didn&#x27;t.

        I think that is less of a factor now, and I think Anthropic have backed off some on being as strict (eg, AFAIK they never implemented the two-tier &quot;claude -p&quot; pricing model they were planning)

      4. spacebanana7 · · focus · HN ↗
        This feels like a horrible precedent. Billing based on data like commits feels like it opens the door to tech stack based billing in general - could we see different prices for people who use other devtools Anthropic doesn&#x27;t like? Makes me feel grateful for open models
      5. saralily · · focus · HN ↗
        Meridian and DirectSDK work well to use a Claude MAX subscription in alternative harnesses.
    16. impulser_ · · focus · HN ↗
      Usage is actually Claude now because of Opus 5.5 since it a better model that Astra. I maxed out my 200$ Claude plan with 10b token on Opus 5 and 5.5 is cheaper. I maxed out two Codex accounts with like not even 5b tokens.
      1. rednb · · focus · HN ↗
        Have you used 10b&#x2F;5b tokens over the course of a week or over the course of a month?
        1. impulser_ · · focus · HN ↗
          It was 9.4B to be exact and it was over the course of two days lol. It was between two projects so 99% of them were cached reads.

          The GPT was about 1B on two projects on 300$ worth of plans all on Astra and I capped out on usage.

          Anthropic caching must be better because the cache rates are better on Claude models.

    17. alansaber · · focus · HN ↗
      It&#x27;s extremely variable because the products are roughly equivelant, and a lot of the quality of service depends on their inference capacity at any given hour&#x2F;day.
    18. huijzer · · focus · HN ↗
      &gt; especially when you factor in ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan.

      I’m currently on the 5x plan and burned through 5% today on a difficult task in 15 minutes so I doubt that. If you got the wrong kind of tasks that you work on, it can go fast.

      1. jeffnash · · focus · HN ↗
        I realized I missed a few words here: I meant &quot;especially when you factor in the fact that ChatGPT usage is unmetered&quot;, i.e. you get unlimited ChatGPT threads that don&#x27;t eat into your codex limit
      2. csnweb · · focus · HN ↗
        But did you use ChatGPT chat or the work mode? Only the former is unmetered at least for me as well.
    19. Marciplan · · focus · HN ↗
      cool! for me its company ethics
      1. platinumrad · · focus · HN ↗
        I don&#x27;t think either of these companies are great, then, but Anthropic is surely worse. The doom marketing is one of the most unethical things an AI company can be doing.
        1. felixgallo · · focus · HN ↗
          tell that nonsense to Huggingface.
          1. platinumrad · · focus · HN ↗
            OpenAI and Anthropic are neck and neck: <a href="https:&#x2F;&#x2F;www.felonybench.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.felonybench.com&#x2F;
            1. felixgallo · · focus · HN ↗
              That&#x27;s a nonsensical website. A third party was responsible for both OpenAI&#x27;s and Anthropic&#x27;s model escapes. The difference is that Anthropic was warning that this might happen, and OpenAI was putting their head in the sand.
              1. platinumrad · · focus · HN ↗
                Your entire comment history really is just this, huh.
    20. TomGarden · · focus · HN ↗
      OpenAI have seemed compute-constrained recently, leading to their subscriptions actually being less generous than Claude as of late. OpenAI even paused purchases of 20x plans.
    21. ChickeNES · · focus · HN ↗
      &gt; that ChatGPT usage (even 6 Astra Pro) is essentially unmetered on the 20x plan

      LMAO, I wish this were true, I hit limits (and the &quot;we are disabling access to protect your data&quot; warnings) all the time, or have chats just...fuck off and get into weird&#x2F;invalid states (interrupted chats, chats that are spinning and stuck, returning &quot;&#x2F;mnt&#x2F;&quot; paths instead of images&#x2F;md files, file links being returned with no file backing them, image classifier firing...and then returning the image anyway (though now I know that GPT-Image-X really really wants to generate NSFW even when that isn&#x27;t the request)).

      Though I am probably an outlier, I have both 20x Claude&#x2F;ChatGPT plans and max both out every week, so... (in my defense I am a hobbyist and this is out-of-pocket)

    22. chrisweekly · · focus · HN ↗
      &gt; &quot;Codex&#x27;s compaction is very good, fwiw, but it happens so frequently that...&quot;

      I appreciate and follow Matt Pocock&#x27;s advice: avoid autocompaction. Compaction is lossy, which is ok when you&#x27;re managing it at phase boundaries, but autocompact is lossy at the most inopportune times, firing mid-task and leading to agents going off the rails.

      1. erichocean · · focus · HN ↗
        Bad advice, compaction is why Codex is so fantastic.

        My conversations compact hundreds of times. By the time it has done a dozen or so compactions, it fully understands the work I want it to do (and how). It&#x27;s almost like having a fine-tuned Astra model.

        10&#x2F;10, would recommend.

        1. chrisweekly · · focus · HN ↗
          I&#x27;m not sure I follow; how is autocompaction (lossy summarization), applied at random times (vs strategically, between workflow phases), helpful to ensuring clarity of intent? Maybe you&#x27;re saying that just plowing ahead and living with the signal loss along the way works well enough for your purposes. In which case, ok, YMMV, different strokes.... but paying attention to context quality and being deliberate about when to compact vs handoff vs delegate to subagents is most definitely not &quot;bad advice&quot;.
          1. edg5000 · · focus · HN ↗
            I agree with @erichocean on this. In theory, compaction is bad. But in practice I found the model is smart enough to write critical details down somewhere, and post-compaction the model doesn&#x27;t make assumptions. A small amount of time is lost reading materials, but the benefit is that you can operate unbounded vs doing small controlled chunks, which is what I used to do with Opus back in the day. Now I just give it as big a task as I can think of.
            1. slopinthebag · · focus · HN ↗
              not my experience at all. compaction during a task is fatal since you lose all of the details of edits and progress halfway through. compacting after task completion is fine though.
            2. elcritch · · focus · HN ↗
              This approach got good with Sol. With 5.5 I&#x27;d break tasks up, record planning docs, etc.

              Now with Sol I rarely bother. It&#x27;s really good at remembering the salient details. Its also great at continuing a pattern I setup, like commit after finishing each feature block, etc.

    23. theshrike79 · · focus · HN ↗
      [delayed]
    24. [deleted] · · focus · HN ↗

      [deleted]

    25. amluto · · focus · HN ↗
      &gt; Codex&#x27;s compaction is very good

      I think it’s only very good in comparison to some of the utter crap that came before it.

      Today I had Codex compaction trigger after I had given an instruction but before it acted on the instruction, and the instruction just disappeared completely. The agent reported that the task was done without actually doing it.

    26. pyinstallwoes · · focus · HN ↗
      What’s this mcp oracle thing?
  38. GodelNumbering · · focus · HN ↗
    Gpt 6 Luna is cheaper than Deepseek 4.1 flash! Today is wild in terms of intelligence&#x2F;price across the board!
    1. gizmodo59 · · focus · HN ↗
      it was expected no? if cost is the only reason to use oss models, they can do much better than small providers who don&#x27;t have much compute.
  39. Someone1234 · · focus · HN ↗
    Have they solved GPT5.6 SOL&#x27;s propensity to over-engineer and over-complicate? You&#x27;d ask SOL to do something relatively simple, and find four single-use methods, an interface, and a factory-factory.

    I actually preferred 5.6-Terra not because it is technically superior (it isn&#x27;t) but because it had better instincts to NOT do this stuff.

    PS - Speaking of better instincts, have they closed the UI-design gap at all? I keep a Claude subscription just because &#x2F;design produces significantly higher quality UI design&#x2F;UI feedback&#x2F;UI refinement than anything I&#x27;ve seen from OpenAI.

    1. cmrdporcupine · · focus · HN ↗
      Astra 6 was a huge improvement over Sol 5.6 for UI work. I haven&#x27;t tried Sol 6 yet for it (it&#x27;s only been a few minutes).

      The GPT &#x2F; Codex models have always been &quot;overengineer&quot; personalities. I prefer that to &quot;I left a pile of race conditions lying around and big gaps in testing&quot; though, which is what I was getting from Opus at times.

      But yes both Astra and Sol veer on the side of paranoid. And honestly that&#x27;s better for team work. For solo work where you just want to yeet something, it can be tiring.

      You learn to tame the GPT &quot;personality&quot; on this front by combing over once a week and asking it to find and exterminate pointless tests, clean abstractions etc.

    2. faitswulff · · focus · HN ↗
      The UI design gap is something I’ve noticed as well, in things as simple as ASCII diagrams. Claude has a more human touch. All the diagrams GPT 5.6 generated for me were dressed up lists with too many pipe symbols.
    3. c0rruptbytes · · focus · HN ↗
      have you tried using lower efforts?
      1. superfrank · · focus · HN ↗
        Not the person you&#x27;re responding to, but I have the same feelings they do and to answer your question for me at least, yes.

        IMO 5.6 Sol had this weird dead zone between medium and high where medium under engineered and took short cuts and high over engineered and ignored instructions it didn&#x27;t agree under the guise of trying being helpful. The whole 5.6 line was the first release from OpenAI where it felt like reasoning level really mattered and was incredibly finicky.

        I haven&#x27;t felt similar issues with GPT 6 though and am very happy with Astra low&#x2F;med&#x2F;high as my default choices depending on the task.

        In general, I felt like with 5.6 the effort level did less than previous to make the models smarter and more just increased the complexity of the response. I have a half joke theory based only on vibes that OpenAI splitting 5.6 into Sol&#x2F;Terra&#x2F;Luna is where the intelligence split happened and so the effort levels were just like &quot;think harder about the decision you already made&quot;. So like if the model decided the earth was flat on low effort it&#x27;d just say something like &quot;the earth is flat because the horizon is flat&quot;. If it was on xhigh reasoning it&#x27;d give you a massively complex answer about how the sun reflects light because of the ozone layer and why people flying in planes can see a curve. In both cases though, adding more effort wouldn&#x27;t get it to realize the earth was round. It just made the answer about it being flat more complex.

        To be clear, that theory is not meant to be taken too seriously. It&#x27;s not based on anything other than vibes. It&#x27;s just my way of explaining to myself something I&#x27;m frustrated about to myself.

        1. mike_hearn · · focus · HN ↗
          Also try just using Luna. It&#x27;s a very capable coding model and doesn&#x27;t over-engineer.
          1. superfrank · · focus · HN ↗
            I&#x27;ve tried 5.6 Luna many times. I don&#x27;t think it&#x27;s any better. I definitely use it from certain tasks, but I find it the most susceptible to that conspiracy theory example I gave above.

            I didn&#x27;t love any of the 5.6 models, but weirdly I think I liked Terra the best. I still wouldn&#x27;t call it amazing though. I&#x27;m still very happy with my codex plan, but 5.6 just wasn&#x27;t my cup of tea I guess.

            Definitely giving 6 Luna and Sol a try this week though.

    4. howunfortunate · · focus · HN ↗
      &gt; design

      I force OpenAI models to use image generation for design, then an iteration loop until it matches the image gen.

      This is frustratingly manual and takes many more repetitions compared to Claude (and especially Claude Design) which &quot;just work&quot;, but it&#x27;s a big step change over the default.

      1. kairosisme · · focus · HN ↗
        FWIW, Codex&#x27;s &quot;Product Design&quot; plugin basically is this workflow (minus built-in iteration, but the model in one prompt will still do its own internal iteration), it&#x27;ll generate 3 images for you to choose from and then build from that + feedback
      2. MisterMunchkin · · focus · HN ↗
        Claude has a bunch of designs hardcoded into it, which is why all of the websites and presentations it makes look the same.
  40. msp26 · · focus · HN ↗
    This Luna pricing is obscene man. 5.6 was good enough for so many use cases (data analysis, structured extraction etc).

    Incredible.

  41. ggcr · · focus · HN ↗
    Live notification in Codex:

    &gt; GPT-5.6-Sol is retiring. This conversation will automatically switch to GPT-6-Sol

    I don&#x27;t recall OAI retiring a model so early lol. Similar arch?

  42. scrlk · · focus · HN ↗
    Artificial Analysis is reporting that 6 Luna scores 2 points lower on their coding index than 5.6 Luna, but is 60% cheaper:

    &gt; In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI&#x27;s Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task.

    <a href="https:&#x2F;&#x2F;x.com&#x2F;ArtificialAnlys&#x2F;status&#x2F;2102462962758033624" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;ArtificialAnlys&#x2F;status&#x2F;2102462962758033624

    Given that they had to discontinue sales of the 20x Pro plan after the Astra release due to compute constraints, I wonder if 6 Sol &amp; Luna are smaller vs their 5.6 counterparts?

    1. Readerium · · focus · HN ↗
      Yup more like a 5.7 than a 6
      1. scrlk · · focus · HN ↗
        Given that they had to discontinue sales of the 20x Pro plan after the Astra launch, I think they&#x27;ve phoned it in with this release as they&#x27;re compute constrained.
    2. scrollop · · focus · HN ↗
      Can we trust AA anymore after the last debacle a week or two ago?
      1. 6thbit · · focus · HN ↗
        wait what debacle?
        1. wyre · · focus · HN ↗
          Probably referencing how when Astra came out it was only 1 point ahead of 5.6 sol.
      2. anthonyrstevens · · focus · HN ↗
        Why is everything a &quot;debacle&quot;. And people complain about Claudisms. sigh
  43. dmitrygr · · focus · HN ↗
    Selling dollar bills for $0.40 to undercut the guys selling them for $0.50 is a bold move. Let&#x27;s see if it pays off for them.
  44. ghoshbishakh · · focus · HN ↗
    So opus 5.5 has reduced price. Who is winning then?
  45. hamburglar1 · · focus · HN ↗
    Code deception 10% at 5.6 to 1.3% for 6.0? So models are getting more safe rather than less safe? hmmm
  46. farceSpherule · · focus · HN ↗

    [dead]

  47. OutOfHere · · focus · HN ↗
    As a user of 5.6-Terra, I am sick and tired of the inconsistencies in GPT model families. There is no 6-Terra.
    1. seizethecheese · · focus · HN ↗
      They dropped Terra because it was worse than Luna &#x2F; Sol at every point of cost performance curve.
      1. OutOfHere · · focus · HN ↗
        That is a fair-sounding but actually invalid argument for multiple reasons:

        1. OpenAI fully controls the user cost for a model, and can set it to where it sits well on the curve.

        2. Performance of shrunken models like Sol&#x2F;Terra&#x2F;Luna is derived from the level of shrinking (relative to Astra). As such, the size and performance of the model is something that is actively targeted when developing the model. If the performance target for Terra was inappropriate for v5.6, this is no way means that it had to be this way for v6.

    2. CharlieDigital · · focus · HN ↗
      I can already see it. 7-Nebula, 8-Galactic, 9-Cosmos; The size inflation is real.
    3. Readerium · · focus · HN ↗
      6 Sol is the new Terra with 50 percent discount vs 5.6 Sol.

      Also 6 Astra Mini would be out soon which would be 5.6 Sol pricing?

      1. OutOfHere · · focus · HN ↗
        I am curious -- what is the basis for claiming that Astra 6 Mini will be out? It doesn&#x27;t make sense to me. It would only further antagonize and confuse users who seek some predictability from a model family.
  48. m3kw9 · · focus · HN ↗
    The new default is 6.0 Sol high. Escalate to Astra-medium. If usage is tight go luna6.0-max
  49. seatac76 · · focus · HN ↗
    Would be funny if Google drops Gemini 4 today.
  50. zaik · · focus · HN ↗
    Why is Claude missing on the &quot;Factuality&quot; graph?
  51. simonw · · focus · HN ↗
    GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

    Here&#x27;s GPT-6 Luna pelicans: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F40d129fc140faca378b9c9f4f16c6ec2" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    And GPT-6 Sol: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    Scroll to the bottom for the GPT-6 Sol max one: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2#response-5" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    For comparison, here are the pelicans I got for GPT-6 Astra: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht... - I still like the Astra Max one best.

    Here&#x27;s a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: <a href="https:&#x2F;&#x2F;static.simonwillison.net&#x2F;static&#x2F;2026&#x2F;gpt-6-and-5.6.html" rel="nofollow">https:&#x2F;&#x2F;static.simonwillison.net&#x2F;static&#x2F;2026&#x2F;gpt-6-and-5.6.h...

    The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.

    1. pantsforbirds · · focus · HN ↗
      The sol max looks like it&#x27;s absolutely ripped for some reason
      1. redanddead · · focus · HN ↗
        He’s been biking a lot
        1. 8bitsout · · focus · HN ↗
          he&#x27;s been cycling a lot
          1. redanddead · · focus · HN ↗
            Ayyy now we’re talking
    2. jdw64 · · focus · HN ↗
      Looking at this, AI still has a long way to go. In Sol Max, the pelican&#x27;s legs are missing on one side—how can one side have two pedals and two legs...
      1. loeg · · focus · HN ↗
        And the bicycles have weird dimensions -- extremely slack head tube angle, handlebars in the wrong orientation, etc.
        1. flyinglizard · · focus · HN ↗
          That&#x27;s just foreshadowing the next generation of 32&quot; all-mountain frames.
    3. dmazin · · focus · HN ↗
      &gt; GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

      Is it? It was already too cheap to meter for me. Luna 6 is actually worse on some benchmarks than 5.6. I’d have loved improved performance for 2x the price than ~equal performance for 0.5x the price.

      1. agentcoops · · focus · HN ↗
        I’ve been doing really heavy text analysis work with LLMs where false negatives&#x2F;misses are important to minimize and my god did I hit cost thresholds quickly with 5.6 Luna — it was the first time I felt motivated to seriously work with local open models, even if inference was degraded for the task. Cheaper and much better inference now brings me back to the closed models for better or worse.
      2. onlyrealcuzzo · · focus · HN ↗
        Hopefully Terra 6 slots somewhat nicely into this space.
      3. FusionX · · focus · HN ↗
        5.6 Luna was already discounted at half the price on OpenRouter. Looks like they made it permanent.
      4. user43928 · · focus · HN ↗
        Yes, I am mildly disappointed with these releases.

        I expected a Fable 5 -&gt; Opus 5 situation, where GPT 6 Sol would perform on par with GPT 6 Astra.

        Instead it&#x27;s more like a price cut on GPT 5.6 Sol, and I&#x27;ll have to stick with Astra for my work.

        The only thing I can hope for is that more users switching to the GPT 6 Sol model frees capacity, allowing OpenAI to hand out some usage resets.

        1. zigzag312 · · focus · HN ↗
          Yeah me too. Maybe that place will occupy the Astra Minor model that appeared in Microsoft&#x27;s Azure model config. As Sol and Sonnet are now similarly priced.
          1. user43928 · · focus · HN ↗
            Good point!

            Maybe they are keeping the cheaper Astra alternative back for their Dev Day next week Tuesday.

    4. gizmodo59 · · focus · HN ↗
      6-luna is at the pareto for most of the tasks! I dont know how they make money here but its insane value from a closed source model. I&#x27;d go further and say it makes no sense (privacy, sovereignty etc aside) to use many other models as its not only expensive but also many providers don&#x27;t have that much GPUs to serve at a significant volume. <a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings?view=month#top-models" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings?view=month#top-models 5.6 luna is already the most used model this month.
      1. krat0sprakhar · · focus · HN ↗
        Can&#x27;t agree more. Between 5.6 Luna and Gemini 3.8 flash I&#x27;m so happy for the value I&#x27;m getting for my dollar (subscription pricing not API pricing) :)
        1. jadbox · · focus · HN ↗
          Gemini 3.8 Flash looks like its better than v7 Luna&#x2F;Sol on DeepSWE v1.1 while at $0.75 per million input tokens and $3.75 per million output tokens. Luna is much cheaper, but Flash has nearly Astra&#x27;s performance for under the price of Sol ($2&#x2F;$10).
          1. antupis · · focus · HN ↗
            Flash thinks much more so it’s pretty much line with Sol for performance. That said I like flash coding style much more than OpenAi models.
            1. jeffnash · · focus · HN ↗
              out of curiosity, what type of code&#x2F;language do you usually use flash to write?
              1. spockz · · focus · HN ↗
                [delayed]
                1. timattrn · · focus · HN ↗
                  what harness or plan are you using 3.8 flash with?
                  1. Kostchei · · focus · HN ↗
                    anti-gravity with gemini 3.8 or gtfo
                  2. spockz · · focus · HN ↗
                    [delayed]
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.