‹ BackHN Continuity

Thread

Reverse-engineered Jev-like model

169 points · 24 comments · rochansinha

  1. vrc · · focus · HN ↗
    Out of curiosity and semi unrelated — why do so many of these projects with customized encoder-decoder setups use earlier Qwen versions like 2.5 and 3 and not the smallest 3.5? Purely the few 100m params, or something else in the latter’s arch or pretraining?
    1. FuckButtons · · focus · HN ↗
      My guess would be qwen 2.5 predates linear attention which would be more complex to use.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.