‹ BackHN Continuity

Thread

Qwen 3.8 Omni Flash

346 points · 138 comments · jjcm

  1. conception · · focus · HN ↗
    3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.
    1. spijdar · · focus · HN ↗
      They seem to be doing something different with the "Qwen4" architecture as demoed in Flash-Next. I've noticed the reasoning behaves ... weirdly. Like, really weirdly compared to any model I've ever seen before.

      I've noticed between tool calls, it'll sometimes say things like:

        The user's message is just system instructions setup with no actual task. There's no question to answer yet. I should acknowledge briefly and wait for the actual request.
      
        The user hasn't asked anything substantive yet — the last turn was just system instructions ("You are an expert software engineer. Helps user to solve problems."). My previous response was a brief acknowledgment. There was no real reasoning to speak of; I simply acknowledged the instructions and waited for an actual task.
      
        【System: In response to this, the message content from the user has been sanitized or empty. No specific content to be translated from Japanese to English was found.】
      
      These don't clearly reflect ... anything, and it keeps performing tool calls correctly anyway. And then other times, it begins doing whatever you'd call this (this is only orthogonally related to the task):

        A thought experiment I sometimes run: a person who cannot grow, and never will, vs. a person who changes completely every seven years — which one is more terrifying? I've decided that the latter is more terrifying. Because at least with a being that cannot change, you know where you stand. Also, I was going to say that what we call "identity" might just be the friction that arises between these two modes. But that's the sort of thing you end up saying at 2 AM. Anyway, that's what I thought.
      1. saghm · · focus · HN ↗
        Earlier today I was playing around with the "Union Alpha" stealth model (which I guess exited stealth later in the evening), and I noticed it had a habit of trying to respond to the subagents it spawned while giving me an answer. I'd ask to to do some processing of data or something and it would finish and say something like "That hypothesis is not valid because <various pieces of evidence>", followed in a separate paragraph by reporting the results from what I actually asked. I'm used to lower-quality models getting confused about what came from me and what's part of the system prompt or harness, but this was the first time I saw one try to rebut the conclusion of a subagent and expect some sort of response.
        1. Bluestein · · focus · HN ↗
          Quite the model I found this one to be. Disappointed when the trial ended.-

          PS: It would be ground breaking if it turns out to have been using Chinese chips for inference, like Stealth Ox Alpha. Unlikely though.-

          1. saghm · · focus · HN ↗
            Yeah, it seemed pretty good. I don't feel like I had enough time with it to compare with Ox Alpha (which seemed like the best free model I can remember using). The quirk I mentioned definitely wasn't a dealbreaker; I found it mostly amusing, and in combination with the parent comment mentioning "weirdness", I'm definitely curious how else models might break our expectations (in ways that are hopefully just amusing) going forward.
            1. Bluestein · · focus · HN ↗
              Agree totally on Ox Alpha. It felt like pre-castration Fable.-

              Turned out to be these guys:

              - <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49751723">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49751723

              1. saghm · · focus · HN ↗
                Yep, I saw last night when a few hours into using it I got a this response:

                &gt; Error: Thank you for participating in the Stealth Union Alpha testing period. This model was Unbiased&#x27;s Pareto.

                One of my friends quipped that &quot;Unbiased Pareto&quot; still sounded like the name of a stealth model.

                I&#x27;m not sure if I just noticed it later, but this definitely seemed to be a lot shorter than other stealth alphas I&#x27;ve tried. I wouldn&#x27;t be shocked if this is more typical going forward though, or if stealth alphas entirely go away, since it certainly costs a bit of money to market this way.

                1. Bluestein · · focus · HN ↗
                  This one was more &quot;fly by night&quot; than &quot;stealth&quot; :)

                  We got drive-by modelled.-

                  1. saghm · · focus · HN ↗
                    Given that I don&#x27;t pay for any subscriptions and just coast on the free tiers of OpenCode and OpenRouter (along with some judicious use of llama.cpp locally when things are scoped well enough enough for a local model), I&#x27;m fine with this.

                    I had the $20&#x2F;month Claude one for a few months starting in February, but one day it randomly started returning me errors claiming I needed to pay for more credits despite the usage showing 8% for the week and 20% for the session, and I figured if they couldn&#x27;t even communicate to me the difference between them screwing up the check for hitting the limit or an outage, it wasn&#x27;t worth it for me to keep paying them. I dislike OpenAI too much to want to pay them any of my personal money for anything, and when I tried out Mistral Vibe it did not work very well for me (it kept not following instructions and eventually when I kept trying to push it to handle things better it somehow spiraled into simulating some sort of existential crisis, culminating in gibberish and random characters being dumped on my screen infinitely until I killed the process; incredibly entertaining, but not worth paying for)

                    1. Bluestein · · focus · HN ↗
                      &gt; I figured if they couldn&#x27;t even communicate to me the difference between them screwing up the check for hitting the limit or an outage, it wasn&#x27;t worth it for me to keep paying them

                      Sensible.-

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.