‹ BackHN Continuity

Thread

Claude Opus 5.5

1806 points · 1134 comments · km144

  1. techjamie · · focus · HN ↗
    With the performance gains they're claiming, I wonder if they implemented the Casual Encoder-Decoder technology from DeepSeek 4.1's paper.

    I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.

    How it works: <a href="https:&#x2F;&#x2F;miraflow.ai&#x2F;blog&#x2F;deepseek-v4-1-flash-causal-encoder-decoder-explained-2026" rel="nofollow">https:&#x2F;&#x2F;miraflow.ai&#x2F;blog&#x2F;deepseek-v4-1-flash-causal-encoder-...

    1. stri8ted · · focus · HN ↗
      This model was likely trained months before deepseek released their paper.
      1. dannyw · · focus · HN ↗
        Models are generally posttrained to a window shorter than you think.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.