‹ BackHN Continuity

Thread

Qwen 3.8 Omni Flash

346 points · 138 comments · jjcm

  1. conception · · focus · HN ↗
    3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.
    1. rubslopes · · focus · HN ↗
      > RL it to oblivion.

      What would that mean in this context?

      1. khafra · · focus · HN ↗
        Others have given examples, but here&#x27;s the theory: <a href="https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;fuSaKr6t6Zuh6GKaQ&#x2F;when-is-goodhart-catastrophic" rel="nofollow">https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;fuSaKr6t6Zuh6GKaQ&#x2F;when-is-go...

        Reinforcement Learning (in LLMs) trains via gradient descent on a reward signal that&#x27;s an imperfect proxy for the actual goal of the engineers doing the training. So, under mild optimization pressure, you get increasingly more of what you want, because that&#x27;s the easiest way to increase the metric.

        But as the optimization pressure increases, so do the ways to increase the metric by doing increasingly weird things. If the full action space grows sufficiently faster than the &quot;things you actually want&quot; subset, the amount of &quot;things you actually want&quot; goes to 0 under sufficient RL.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.