‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. lwansbrough · · focus · HN ↗
    Anyone else more excited about Chinese models than American models these days? Big thing for me is affordability.
    1. joshheitzman · · focus · HN ↗
      Absolutely! DeepSeek-V4-Flash-0731 has become my daily driver. It's pretty amazing what it can do for what it costs at deepinfra.com (I don't use deepseek as a provider since they train on your data [at least their honest about it]). GLM-5.1 was my daily driver before that and Kimi K2.5 before that.
      1. kingforaday · · focus · HN ↗
        Are you finding DS better then kimi k3 and glm-5.3? Do you mind sharing your primary use case?
        1. joshheitzman · · focus · HN ↗
          My primary use is AI coding agent. Its vastly cheaper than Kimi K3 and I haven&#x27;t found a scenario where I really need Kimi K3 versus smaller models. GLM-5.3 Flash is good but there is series of bugs in the vllm middleware that prevent GLM models from getting all of their reasoning content returned to them that impairs inference quality. A lot of inference providers use vllm which makes it hard to find a good provider for GLM. I&#x27;ve been using friendli.ai but using GLM-5.3 Flash from them is more expensive then using DS V4 Flash from deepinfra.com simply because deepinfra.com is so cheap. The DS V4 Flash cost at together.ai is similar to the GLM-5.3 Flash from friendli.ai or at least that&#x27;s what I found in my benchmarks a week ago: <a href="https:&#x2F;&#x2F;www.linkedin.com&#x2F;posts&#x2F;joshheitzman_i-ran-a-fuller-round-of-benchmarks-on-my-activity-7506040800166723584-WN3v" rel="nofollow">https:&#x2F;&#x2F;www.linkedin.com&#x2F;posts&#x2F;joshheitzman_i-ran-a-fuller-r...
        2. pimeys · · focus · HN ↗
          I&#x27;ve used Kimi K3 for a few months as my main model and DeepSeek 4.1 is as fast and about 10x cheaper.

          I just had like four big sessions going today, paid about $8 in tokens. I see no reason to pay more, this is more than I need for intelligence.

          1. pkulak · · focus · HN ↗
            4.1 consistently surprises me in capability for the price. And I don&#x27;t think I&#x27;m the only one. It&#x27;s been dominating the leaderboard at OpenRouter, and I just got an email today from Fireworks saying they were _raising_ the price by about 30%. I&#x27;ll probably switch, because their infra doesn&#x27;t support being the highest-cost, but it&#x27;s still telling.
            1. celrod · · focus · HN ↗
              I tried it a few times and liked the speed, but often found it ended up looping, i.e. repeating the same token sequence (e.g. the same sequence of 5 paragraphs) over and over again until it hit the max output limit. This doesn&#x27;t end up happening every session, but does every now and then.

              My impression of DSv4.1-flash was very positive aside from this. But that was enough for me to stick with GLM-5.3(-flash), which both gave me consistently great results

              I was using a vibe coded bare bones harness. I was wondering if this was normal from DSv4.1-flash, or if its my harnesses fault.

              1. pkulak · · focus · HN ↗
                I&#x27;ve had that looping issue with open models too. But never 4.1. I wonder if it&#x27;s a model + harness combo? But yeah, one loop issue and I&#x27;m done with a model forever.
                1. pimeys · · focus · HN ↗
                  Harness. Especially if a tool call error doesn&#x27;t say what to do next and the model is not RL&#x27;d with that tool, a retry storm is common.

                  So if you use MCP a lot, simplify the params, be more lenient on validation and rework the errors.

                  It is quite good with shell.

                  1. celrod · · focus · HN ↗
                    Yeah, that&#x27;s what I&#x27;d been leaning towards. No mcp, but I&#x27;ll see if I can reproduce and debug it, since other people don&#x27;t seem to have that problem as badly as I&#x27;ve experienced it (and the idea of having a nasty bug like that bothers me).

                    No mcp support. I&#x27;ll try copying deepseek harness&#x27;s basic tool call formats as a starting point.

      2. tristanMatthias · · focus · HN ↗
        How does it compare to 4.1 flash? Curious why folks don’t use the more “modern” one.
        1. joshheitzman · · focus · HN ↗
          I haven&#x27;t tried 4.1 flash as I&#x27;m assuming its a preview. I did not get good results from the preview version of 4.0 flash (i.e. the one that did not include the month and date of release in its name).
          1. CamperBob2 · · focus · HN ↗
            4.1 Flash is a horse of a very different color. It cooks. IMHO it&#x27;s probably a preview of DS5, rather than a true DS4-series model.
        2. randbyte · · focus · HN ↗
          4.1 flash is very fast and capable. Token efficiency is not great so it fill up context window much faster compared to similarly capable models.

          glm 5.3 flash is a tad slower but a bit more capable and way more token efficient.

          Source: self hosted tested on rented GB200 node at 8bit.

          1. pkulak · · focus · HN ↗
            Wow, I&#x27;m surprised you are saying GLM 5.3 Flash is more capable. Isn&#x27;t is like half the price of 4.1 Flash?
            1. randbyte · · focus · HN ↗
              I don’t know. They are self hosted so I am not comparing token cost.

              ds 4.1 is lightning fast though. Also much better in image recognition.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.