‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

816 points · 361 comments · snehesht

  1. Jackson__ · · focus · HN ↗
    I've just tested Strata on a simple 50 image vision benchmark. The task is to output the exact coordinates of a requested object. The result via Strata had a median error distance of 154.8 pixels, avg of 168.8. Running the exact same GGUF and vision adapter weights on llama.cpp gives me a median error of 46.5, avg 81.4.

    To put that into perspective, here are some more numbers from other models via llama.cpp:

    Median/Average

    Qwen 3.5 9B BF16: 46.5 / 193.3

    Qwen 3.6 35B Q4 K XL: 38.4 / 76.4

    Qwen 3.5 122B Q3 K M: 32.9 / 68.6

    The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.

    I have done no further testing, as these results line up perfectly with my expectations.

    1. NamlchakKhandro · · focus · HN ↗
      Tldr, strata is a waste of time.
      1. unlikelytomato · · focus · HN ↗
        at least for vision? Are there similar comparisons for language? It seems like vision is often an afterthought when it comes to bootstrapping these newer inference engines
        1. Borealid · · focus · HN ↗
          Vision functions the same way as language when inference is done. It's a stream of tokens.

          It's just easier to measure the "right" answer (and deviation therefrom) on a vision task than a language one due to the underspecified nature of language.

          1. pjc50 · · focus · HN ↗
            How does image tokenizing work?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.