‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. MisterMunchkin · · focus · HN ↗
    It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.

    And my job won’t even pay for Claude now because it’s so ruinously expensive.

    1. yipinwong · · focus · HN ↗
      Say that to Luna's face. Ya all bringing up this not-so-cheap-nowadays chinese models and not that more intelligent than luna and bringing "cost" as the only factor.
      1. cromka · · focus · HN ↗
        Not even GPT6 Sol cannot match DeepSeek 4.1 in my work, with outrageous bugs. Don't get me started on Luna.

        From my today's session with Sol:

        - it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.

        - for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.

        - it said it would ask me to approve/amend the suggested PR message, it never did and fired off right away

        - it keeps forgetting the changes it did itself; no context compaction was used

        - it said it tested the change visually, but it did not even try

        - hallucinated several facts despite me asking beforehand to check online.

        On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.

        This is astonishingly bad and it is nowhere close to what Sol 5.6 was a month ago.

        I can't deal with this sh*t anymore, I have no trust in the tools I use and both OpenAI and Anthropic do the same thing.

        1. gruez · · focus · HN ↗
          >Not even GPT6 Sol cannot match DeepSeek 4.1 in my work, with outrageous bugs.

          That seems hard to believe even with deepseek's own benchmarks. Not to mention for every person who says chinese ai is ahead of american labs, there's like 10 saying that they're benchmaxxed or that they're merely "decent value for money".

          1. MitziMoto · · focus · HN ↗
            Yeah, I don't know what model that guy is using. He's talking about Sol like it's GPT4.

            I haven't had these issues in multiple model generations of models.

            Failing to understand English grammar? Give me a break.

            1. NewsaHackO · · focus · HN ↗
              You can tell how vague the grammar complaint was it probably wasnt a big of a gaffe that they think it was. Also, slight tangent, but the part where he said that it admitted where there was an error but was not able to say why it made the error is one of the issues with model sycophancy; it want to validate the users feelings of being wronged, but also does not want to say something factually incorrect. So it produces this fail state of saying sorry for nothing, but it cannot backtrack.
            2. cromka · · focus · HN ↗
              > me: "The tool divides network rates by 1,000, but uses the \`KB\` instead of \`kB\`. \`KB\` (alongside the SI-standardized KiB), is reserved for units based on 1,024. This change corrects the usage of those units.

              > it: That still attributes two claims to SI that SI does not make. `KB` is not reserved for 1,024 bytes, and `KiB` is the standardized binary symbol.

              > me: where does it say that SI reserves KB?

              > it: Nowhere. You said *KiB* was SI-standardized; you did not say SI reserves *KB*. I misread your sentence and argued against a claim you did not make.

              It was correct to point me out on my mistake in essence, but still misunderstood my bracketed "(alongside the SI-standardized KiB)" sentence.

              Sure, it wasn't a grammar mistake as such, more like a logical one, but it still shouldn't make it. I had more than one such issues already with it, this one was most pronounced.

              What it interesting, though, is the number of corporate apologists my comment brought in. It's like it doesn't matter how many times OpenAI and Anthropic have botched some of the models while keeping the branding, some people would still die on that apologist hill.

          2. solenoid0937 · · focus · HN ↗
            HN desperately wants Chinese AI to be competitive so there's a lot of self delusion and wishful thinking going on.

            I use Deepseek 4.1 almost every day as well, it's nowhere close

            1. stymaar · · focus · HN ↗
              There's an interesting contradiction in your comment: either the Chinese models aren't competitive, or you wouldn't be using it!
              1. mannycalavera42 · · focus · HN ↗
                maybe it's the combination of both? the _could_ have a judgment backed by firsthand experience for example...woah, I know!

                It's not like rooting for a favorite football team. double woah!

                1. stymaar · · focus · HN ↗
                  Why would they use it if it's not competitive in some way though?
      2. pavo-etc · · focus · HN ↗
        Just yesterday I ran a comparison of $/message through my harness[0] comparing Deepseek models to Luna, and to my surprise Luna won. I suspect its partially due to Deepseek's long thinking times, and also maybe due to OpenRouter variance in cache pricing etc.

        Model and observed window | Messages | Retrieved actual $/message

        DeepSeek v4-flash — all observed snapshots, 20 Jul–15 Sep | 4,610 | $0.0241

        DeepSeek v4.1-flash — 11–27 Sep, before the 28 Sep billing change | 1,079 | $0.0576

        DeepSeek v4.1-flash — 28 Sep, partial new billing window | 69 | $0.0365

        GPT-6 Luna — 23–28 Sep, partial final day | 88 | $0.0329

        I've subbed to Codex because I suspect at my usage rates the Codex Plus plan gives me more Luna messages than I'm using, and I've not really observed and better or worse intelligence performance. Interested to see how my $/message comes out after a month of usage on the Codex plan.

        Something nice I've realised about my harness is that I can run different agents on different models so I can collect pricing data for a bunch in parallel.

        [0]: pi-msg, run pi agents over xmpp <a href="https:&#x2F;&#x2F;github.com&#x2F;zachpmanson&#x2F;pi-msg" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;zachpmanson&#x2F;pi-msg

      3. stymaar · · focus · HN ↗
        Why would you use a cloud model that&#x27;s no better than a local one though? Because if you pick Luna then you don&#x27;t have to compare it to the big Chinese models, Qwen3.8-27B is what you want to compare it to, and the comparison doesn&#x27;t make Luna look good.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.