‹ BackHN Continuity

Thread

Qwen3.8 Max now ranked as the best overall model by agentic index

252 points · 127 comments · apitman

  1. syntaxing · · focus · HN ↗
    I am so excited for Qwen 3.8 27B. It’s a shame how slow prefill (~3-400) is on a strix halo but it’s such a good model for agentic tasks.
    1. tarr11 · · focus · HN ↗
      What type of agentic tasks are you using it for (eg how complex)?
      1. syntaxing · · focus · HN ↗
        For personal stuff, I use it with AnythingLLM. It replaced any Google search for me. For coding, I run opencode though I have been debating switching to Pi.
    2. CamperBob2 · · focus · HN ↗
      How are you running it on a Strix Halo? The weights aren't out yet, are they?
      1. 13rac1 · · focus · HN ↗
        I interpret @syntaxing as meaning they are looking forward to running Qwen3.8-27B, but are frustrated by prefill times with other models, such as Qwen3.6-27B.
      2. syntaxing · · focus · HN ↗
        I meant Qwen3.6. Unsloth supposedly has early preview of the model and the VRAM requirement is the same so most people expect similar model size and type.
    3. LoganDark · · focus · HN ↗
      I find that 35B-A3B is much easier to run on my M4 Max (both prefill and generation)
      1. markasoftware · · focus · HN ↗
        It's well known 35b is much faster (on any hardware) and quite a bit dumber
    4. colingauvin · · focus · HN ↗
      Prefill is survivable if you cache well. But what kills me is the context. Qwen 27 needs a ton of room for KV Cache. I guess not an issue on a 128 GB Halo or Spark, but if you are running of consumer/prosumer GPUs it's miserable to be compacting every 120k tokens.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.