‹ BackHN Continuity

Thread

Claude Opus 5.5

1806 points · 1134 comments · km144

  1. techjamie · · focus · HN ↗
    With the performance gains they're claiming, I wonder if they implemented the Casual Encoder-Decoder technology from DeepSeek 4.1's paper.

    I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.

    How it works: <a href="https:&#x2F;&#x2F;miraflow.ai&#x2F;blog&#x2F;deepseek-v4-1-flash-causal-encoder-decoder-explained-2026" rel="nofollow">https:&#x2F;&#x2F;miraflow.ai&#x2F;blog&#x2F;deepseek-v4-1-flash-causal-encoder-...

    1. ryangg · · focus · HN ↗
      Getting a 403 on that link. Mind checking it once?
      1. potwinkle · · focus · HN ↗
        I&#x27;m able to access it on my laptop at home. Maybe a misconfigured bot protection rule, try a different user-agent or IP?
      2. zatkin · · focus · HN ↗
        It&#x27;s working for me (based out of California).
      3. peri-cl · · focus · HN ↗
        <a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260922172456&#x2F;https:&#x2F;&#x2F;miraflow.ai&#x2F;blog&#x2F;deepseek-v4-1-flash-causal-encoder-decoder-explained-2026" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260922172456&#x2F;https:&#x2F;&#x2F;miraflow....

        tired: AI startup attempting to build their own website

        wired: a nonprofit founded in 1996

    2. stri8ted · · focus · HN ↗
      This model was likely trained months before deepseek released their paper.
      1. manquer · · focus · HN ↗
        Doesn&#x27;t mean they didn&#x27;t apply something similar. They could have also come up independently with their own version, the speculation is not they copied it, rather that they have performance breakthroughs which perhaps is a result of work in same domain
      2. dannyw · · focus · HN ↗
        Models are generally posttrained to a window shorter than you think.
    3. ACCount39 · · focus · HN ↗
      I don&#x27;t think it&#x27;s particularly relevant?

      They might be using something like this, or they might be using some other &quot;increased sparsity&quot; techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.

      Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that&#x27;s unlikely though.

    4. Balinares · · focus · HN ↗
      Unless they already have something similar of their own, which is always possible, they&#x27;d be stupid not to. I don&#x27;t suppose we&#x27;ll ever know, though. It would not be a good look if after the trillions of dollars that have been thrown at US labs, investors found out that they&#x27;re down to copying Chinese tech.
    5. MisterMunchkin · · focus · HN ↗
      That post is just AI slop. You forgot to strip out the slop lines.
    6. cpldcpu · · focus · HN ↗
      You mean the one they copied from Microsofts paper? (properly cited as well)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.