‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. MisterMunchkin · · focus · HN ↗
    It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.

    And my job won’t even pay for Claude now because it’s so ruinously expensive.

    1. edu · · focus · HN ↗
      What model are you using ?
      1. system2 · · focus · HN ↗
        Not him but 3 models dominate: GLM 5.3, Qwen 3.8, Mimo 2.6. All censoring certain things. Numbers and other uses are perfectly fine. They are like 0.10-0.15 per 1M tokens. American AI lost the game already, people just can't see it.
        1. case540 · · focus · HN ↗
          People dont want iPhone 14s in 2026. People want the latest and greatest. Chinese companies desperately trying to get western usage of their models
          1. nehal3m · · focus · HN ↗
            They would if those phones were 899/10=89.9 bucks.
            1. keyworkorange · · focus · HN ↗
              How are these models so cheap?
              1. system2 · · focus · HN ↗
                They are just not overcharging. Nvidia's top AI chip Rubin sells in 72-GPU racks for about $3.5–7.8M. A rack running Xiaomi's MiMo V2.6 Pro generates roughly 150–300B tokens a day, worth about $130–260k at Xiaomi's API price.

                That's a payback of the infrastructure in a few weeks in theory. After a few weeks or a month, the only cost is electricity, and whatever they make after that is pure profit. This is why they can charge normal prices. Not $50 for 1M tokens.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.