‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. MisterMunchkin · · focus · HN ↗
    It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.

    And my job won’t even pay for Claude now because it’s so ruinously expensive.

    1. throwa356262 · · focus · HN ↗
      Obviously not as "intelligent" but almost 10x cheaper

      Mimo 2.6 Pro: 0.04/0.4/0.87

      Sonnet 5.5: 0.2/2/10

      Opus 5.5: Sonnet prices times 2

      What I dont understand is their cache writes ($2.5). Why is that not covered by input cost?

      1. tintor · · focus · HN ↗
        You don't have to pay for cache write if prompt isn't part of conversation.
      2. lcampbell · · focus · HN ↗
        I was under the impression that the cache write fee was added to both the input and output costs (except in cases where the cache write is explicitly disabled via e.g. DISABLE_PROMPT_CACHING). The output becomes part of the context, after all; if they don't (for some reason, due to disaggregated inference perhaps) then I'd expect output tokens get charged both output then input+cache_write on the subsequent completion request.

        The pricing model confuses me though (I presume by design, Hanlon be damned).

      3. [deleted] · · focus · HN ↗

        [deleted]

      4. [deleted] · · focus · HN ↗

        [deleted]

      5. anvuong · · focus · HN ↗
        Which provider are you using Mimo from? The official Xiaomi one seems to be very slow/buggy for me in the last few days.
    2. edu · · focus · HN ↗
      What model are you using ?
      1. system2 · · focus · HN ↗
        Not him but 3 models dominate: GLM 5.3, Qwen 3.8, Mimo 2.6. All censoring certain things. Numbers and other uses are perfectly fine. They are like 0.10-0.15 per 1M tokens. American AI lost the game already, people just can't see it.
        1. zarmin · · focus · HN ↗
          What harness are you using with those?
          1. beveradb · · focus · HN ↗
            opencode
        2. toasty228 · · focus · HN ↗
          > American AI lost the game already, people just can't see it.

          Microsoft has been releasing dog shit insanely overpriced software with decent alternatives for decades and is still used in every single company I work for or with.

          Your take is the "current year is the year of the linux desktop" meme of "ai"

          1. Lord-Jobo · · focus · HN ↗
            The difference is that Microsoft did that while relying on the extremely load bearing windows ecosystem. These AI companies have no equivalent lock in, nothing even close to it honestly.
            1. toasty228 · · focus · HN ↗
              I can guarantee you 80% of people will call any llm "a chatgpt", most have never heard of claude, even less of opus, or sonnet, "deepseek" probably reminds them of a brand of toothpaste or something like that, "GLM" might make them think of the new mercedes SUV perhaps. 99.9% will never self host, nor send a single sent to a chinese model provider.
              1. system2 · · focus · HN ↗
                They are not the ones spending API money.
                1. toasty228 · · focus · HN ↗
                  I don't know a single company using deepseek internally in any capacity, and I have friends in a lot of tech/tech heavy companies, virtually all of them use claude, the lucky ones get cursor with claude/chatgpt/grok.
                  1. cromka · · focus · HN ↗
                    Interestingly, I hear about them doing that all the time. But this is in EU.
                  2. system2 · · focus · HN ↗
                    Once again, just like your other reply, you are mixing things up- a logical fallacy called the straw man; look it up. Your responses usually use this type of misdirection. Coding agent =/= API.

                    AI does not mean coders coding with agents all day long. AI integration is mostly for data processing, which is the promise and use of the APIs.

            2. noisy_boy · · focus · HN ↗
              Most people I know working in banks etc are using Co-pilot. Not because it is any good, because it integrates with the entire MS ecosystem - management pushes it. I would suspect that is the main driver of MS's AI penetration numbers for desktop-level non-API usage.
          2. system2 · · focus · HN ↗
            Most businesses I know switched to Google Sheets or Google Docs.
            1. toasty228 · · focus · HN ↗
              Never heard about anyone using google sheets and docs for messaging, email, presentations, etc.
              1. system2 · · focus · HN ↗

                [dead]

              2. joseda-hg · · focus · HN ↗
                Gmail, Meet and Presentations do all of those

                People use them, if for no other reason, because they are cheap, or are part of the Chromebook generation and have gotten used to it

                Of their suite, Presentation and Sheets are the only ones people really have gripes about, Sheets by power users because it isn't Excel and it can never be, and Presentations because it's the ugly duckling of the suite

              3. conception · · focus · HN ↗
                You’ve never heard of anyone using gmail for email?
              4. madeofpalk · · focus · HN ↗
                Huh. I’ve never worked at a company that doesn’t primarily/exclusively use Google Docs/Sheets/Slides. Guess that goes to show we’re really all in our own little bubble.
            2. presentation · · focus · HN ↗
              Basically only true in the Silicon Valley bubble. Professional services are Microsoft Office/SharePoint all in.
          3. BeetleB · · focus · HN ↗
            You're a decent sized company and wants to manage the SW + security on all your employee's PCs. They need to be able to update/remote SW on your machine remotely, see your settings, etc.

            I don't think anything comes close to Microsoft's offerings. Macs suck. Ditto Linux.

            1. SSLy · · focus · HN ↗
              [delayed]
          4. csomar · · focus · HN ↗
            When Microsoft did that there wasn't really any competition and they were also the cheap option. I don't think they would survive if there was a sufficiently good chinese OS at the time.
          5. fearmerchant · · focus · HN ↗
            It could be a Linux vs Mac/Win situation though. Sure the HN crowd will build their own harnesses and hook up cheaper models but the average Joe Sixpack wants something that "just works" out of box and the frontier labs seems well positioned to give that to them.
        3. tripleee · · focus · HN ↗
          From my experience GLM 5.3 is at most 6 months away from the frontier models, and good enough for most tasks already

          Even 5.2 is doing really well in comparison here: <a href="https:&#x2F;&#x2F;labs.scale.com&#x2F;leaderboard&#x2F;sweatlas-refactoring" rel="nofollow">https:&#x2F;&#x2F;labs.scale.com&#x2F;leaderboard&#x2F;sweatlas-refactoring

        4. swingandamiss · · focus · HN ↗
          Europe didn&#x27;t even start in the race. Europe is the biggest losers of all it seems. What a shame.
          1. etatester · · focus · HN ↗
            The reason why that happens is not really Europe&#x27;s fault. Any remotely-competent developer just works for US companies, why would they not? Especially for AI, there just isn&#x27;t enough investments in the EU to offer those salaries.
        5. case540 · · focus · HN ↗
          People dont want iPhone 14s in 2026. People want the latest and greatest. Chinese companies desperately trying to get western usage of their models
          1. nehal3m · · focus · HN ↗
            They would if those phones were 899&#x2F;10=89.9 bucks.
            1. keyworkorange · · focus · HN ↗
              How are these models so cheap?
              1. system2 · · focus · HN ↗
                They are just not overcharging. Nvidia&#x27;s top AI chip Rubin sells in 72-GPU racks for about $3.5–7.8M. A rack running Xiaomi&#x27;s MiMo V2.6 Pro generates roughly 150–300B tokens a day, worth about $130–260k at Xiaomi&#x27;s API price.

                That&#x27;s a payback of the infrastructure in a few weeks in theory. After a few weeks or a month, the only cost is electricity, and whatever they make after that is pure profit. This is why they can charge normal prices. Not $50 for 1M tokens.

        6. jpgvm · · focus · HN ↗
          What are they censoring that matters to programmers?
        7. anvuong · · focus · HN ↗
          Does AI-style writing start bleeding into the comments, or is HN now also full of bots like reddit?
        8. ndm000 · · focus · HN ↗
          This ignores two things.

          OpenAI and Anthropic have both transitioned into product companies. ChatGPT (the app) and Claude are both one-click installs that just work. People and businesses with pay for this.

          People will also pay for the best (or the perception of being the best). Since it&#x27;s hard to tell what &quot;intelligence&quot; really means model to model, there&#x27;s a sense of safety in giving a task to the &quot;best&quot;.

        9. Art9681 · · focus · HN ↗
          The usage numbers tell a different story. Not only is the great majority of AI users on American frontier providers, they are also willing to pay the premium price for the premium product. It&#x27;s not like tech is oblivious to the Chinese models. It&#x27;s that the industry is aware they are always 6 months to a year behind.

          You didn&#x27;t discover some new trick for cost performance. And the rest of the world isn&#x27;t dumb.

          You&#x27;re just too broke to afford the supercar and justifying the hooptie. It gets you to the kindergarten class after all. And that&#x27;s all you need.

    3. yipinwong · · focus · HN ↗
      Say that to Luna&#x27;s face. Ya all bringing up this not-so-cheap-nowadays chinese models and not that more intelligent than luna and bringing &quot;cost&quot; as the only factor.
      1. cromka · · focus · HN ↗
        Not even GPT6 Sol cannot match DeepSeek 4.1 in my work, with outrageous bugs. Don&#x27;t get me started on Luna.

        From my today&#x27;s session with Sol:

        - it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn&#x27;t explain why it made it.

        - for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.

        - it said it would ask me to approve&#x2F;amend the suggested PR message, it never did and fired off right away

        - it keeps forgetting the changes it did itself; no context compaction was used

        - it said it tested the change visually, but it did not even try

        - hallucinated several facts despite me asking beforehand to check online.

        On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.

        This is astonishingly bad and it is nowhere close to what Sol 5.6 was a month ago.

        I can&#x27;t deal with this sh*t anymore, I have no trust in the tools I use and both OpenAI and Anthropic do the same thing.

        1. gruez · · focus · HN ↗
          &gt;Not even GPT6 Sol cannot match DeepSeek 4.1 in my work, with outrageous bugs.

          That seems hard to believe even with deepseek&#x27;s own benchmarks. Not to mention for every person who says chinese ai is ahead of american labs, there&#x27;s like 10 saying that they&#x27;re benchmaxxed or that they&#x27;re merely &quot;decent value for money&quot;.

          1. MitziMoto · · focus · HN ↗
            Yeah, I don&#x27;t know what model that guy is using. He&#x27;s talking about Sol like it&#x27;s GPT4.

            I haven&#x27;t had these issues in multiple model generations of models.

            Failing to understand English grammar? Give me a break.

            1. NewsaHackO · · focus · HN ↗
              You can tell how vague the grammar complaint was it probably wasnt a big of a gaffe that they think it was. Also, slight tangent, but the part where he said that it admitted where there was an error but was not able to say why it made the error is one of the issues with model sycophancy; it want to validate the users feelings of being wronged, but also does not want to say something factually incorrect. So it produces this fail state of saying sorry for nothing, but it cannot backtrack.
            2. cromka · · focus · HN ↗
              &gt; me: &quot;The tool divides network rates by 1,000, but uses the \`KB\` instead of \`kB\`. \`KB\` (alongside the SI-standardized KiB), is reserved for units based on 1,024. This change corrects the usage of those units.

              &gt; it: That still attributes two claims to SI that SI does not make. `KB` is not reserved for 1,024 bytes, and `KiB` is the standardized binary symbol.

              &gt; me: where does it say that SI reserves KB?

              &gt; it: Nowhere. You said *KiB* was SI-standardized; you did not say SI reserves *KB*. I misread your sentence and argued against a claim you did not make.

              It was correct to point me out on my mistake in essence, but still misunderstood my bracketed &quot;(alongside the SI-standardized KiB)&quot; sentence.

              Sure, it wasn&#x27;t a grammar mistake as such, more like a logical one, but it still shouldn&#x27;t make it. I had more than one such issues already with it, this one was most pronounced.

              What it interesting, though, is the number of corporate apologists my comment brought in. It&#x27;s like it doesn&#x27;t matter how many times OpenAI and Anthropic have botched some of the models while keeping the branding, some people would still die on that apologist hill.

          2. solenoid0937 · · focus · HN ↗
            HN desperately wants Chinese AI to be competitive so there&#x27;s a lot of self delusion and wishful thinking going on.

            I use Deepseek 4.1 almost every day as well, it&#x27;s nowhere close

            1. stymaar · · focus · HN ↗
              There&#x27;s an interesting contradiction in your comment: either the Chinese models aren&#x27;t competitive, or you wouldn&#x27;t be using it!
              1. mannycalavera42 · · focus · HN ↗
                maybe it&#x27;s the combination of both? the _could_ have a judgment backed by firsthand experience for example...woah, I know!

                It&#x27;s not like rooting for a favorite football team. double woah!

                1. stymaar · · focus · HN ↗
                  Why would they use it if it&#x27;s not competitive in some way though?
      2. pavo-etc · · focus · HN ↗
        Just yesterday I ran a comparison of $&#x2F;message through my harness[0] comparing Deepseek models to Luna, and to my surprise Luna won. I suspect its partially due to Deepseek&#x27;s long thinking times, and also maybe due to OpenRouter variance in cache pricing etc.

        Model and observed window | Messages | Retrieved actual $&#x2F;message

        DeepSeek v4-flash — all observed snapshots, 20 Jul–15 Sep | 4,610 | $0.0241

        DeepSeek v4.1-flash — 11–27 Sep, before the 28 Sep billing change | 1,079 | $0.0576

        DeepSeek v4.1-flash — 28 Sep, partial new billing window | 69 | $0.0365

        GPT-6 Luna — 23–28 Sep, partial final day | 88 | $0.0329

        I&#x27;ve subbed to Codex because I suspect at my usage rates the Codex Plus plan gives me more Luna messages than I&#x27;m using, and I&#x27;ve not really observed and better or worse intelligence performance. Interested to see how my $&#x2F;message comes out after a month of usage on the Codex plan.

        Something nice I&#x27;ve realised about my harness is that I can run different agents on different models so I can collect pricing data for a bunch in parallel.

        [0]: pi-msg, run pi agents over xmpp <a href="https:&#x2F;&#x2F;github.com&#x2F;zachpmanson&#x2F;pi-msg" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;zachpmanson&#x2F;pi-msg

      3. stymaar · · focus · HN ↗
        Why would you use a cloud model that&#x27;s no better than a local one though? Because if you pick Luna then you don&#x27;t have to compare it to the big Chinese models, Qwen3.8-27B is what you want to compare it to, and the comparison doesn&#x27;t make Luna look good.
    4. usef- · · focus · HN ↗
      If it&#x27;s for personal use, is there a reason you don&#x27;t want the subscription?

      Anthropic&#x27;s $20 subscription gives &gt;$500 worth of credit by most measures, which is pretty similar, and you get a better model. Their raw API prices have fat margins.

      And as another commenter said, Luna is the cost leader at the moment if you really need API pricing.

      1. daemonologist · · focus · HN ↗
        For personal use I simply need less than $20 of tokens, even at API rates. If they (Anthropic or OpenAI) offered a $5&#x2F;month plan that gave you ~$50 of credit I would probably sign up.
    5. holbrad · · focus · HN ↗
      That&#x27;s only true if you&#x27;re paying API prices, but you really, really should be using the subscriptions. Both the personal and business ones are still really good value.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.