‹ BackHN Continuity

Thread

Ember-1

589 points · 249 comments · gmays

  1. ls612 · · focus · HN ↗
    On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too.
    1. spijdar · · focus · HN ↗
      I suspect it might generalize to other large models, but I don't think Qwen3.8 27B is one of them. Kimi K3 is a 2.8 trillion parameter model, and I suspect that is playing a big role in being able to reduce the length of CoT without taking a hit in quality.

      That's just vibes, though.

    2. KaoruAoiShiho · · focus · HN ↗
      <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1wj3s31&#x2F;thank_you_swift_qwen_38_27b_now_has_100k&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1wj3s31&#x2F;thank_y...
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.