‹ BackHN Continuity

Thread

Reverse-engineered Jev-like model

169 points · 24 comments · rochansinha

  1. vrc · · focus · HN ↗
    Out of curiosity and semi unrelated — why do so many of these projects with customized encoder-decoder setups use earlier Qwen versions like 2.5 and 3 and not the smallest 3.5? Purely the few 100m params, or something else in the latter’s arch or pretraining?
    1. augment_me · · focus · HN ↗
      In my experience if you tell Claude to port LLM-like stuff without explicit steering for versioning, it will default to the most popular thing for this in its training window to reduce errors. 3.5 is outside its training data.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.