‹ BackHN Continuity

Thread

Reverse-engineered Jev-like model

169 points · 24 comments · rochansinha

  1. vrc · · focus · HN ↗
    Out of curiosity and semi unrelated — why do so many of these projects with customized encoder-decoder setups use earlier Qwen versions like 2.5 and 3 and not the smallest 3.5? Purely the few 100m params, or something else in the latter’s arch or pretraining?
    1. brainless · · focus · HN ↗
      I came across this recently. I was scanning for tiny models from HF using their search API. The script was generated by an agent. When I ran it, Qwen 3.5 did not make it at the top. Turns out, models generally prefer older content (training) but that the scanner also did not give any importance to recency.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.