‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

295 points · 177 comments · apitman

  1. aliljet · · focus · HN ↗
    Is there a path to distill this model to do very specific things? Like a RAG strategy for a small (or even large) corpus?
    1. Alpha3031 · · focus · HN ↗
      Depends on what you want to do. Some task specific models can be trained with a few ten or hundred thousand training examples so you can use a bigger model to produce synthetic training examples and then fine tune a smaller student model. I think that's the usual process. Whether you'd get acceptable performance this way depends, as mentioned, on what you're trying to do and what you'd consider acceptable.
    2. teravor · · focus · HN ↗
      once you are able to get the full probability distributions per token you can distill it on specific domains. distilling without that isn't generally a good idea unless you have invested millions in the requisite infrastructure.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.