‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

793 points · 354 comments · snehesht

  1. Jackson__ · · focus · HN ↗
    I've just tested Strata on a simple 50 image vision benchmark. The task is to output the exact coordinates of a requested object. The result via Strata had a median error distance of 154.8 pixels, avg of 168.8. Running the exact same GGUF and vision adapter weights on llama.cpp gives me a median error of 46.5, avg 81.4.

    To put that into perspective, here are some more numbers from other models via llama.cpp:

    Median/Average

    Qwen 3.5 9B BF16: 46.5 / 193.3

    Qwen 3.6 35B Q4 K XL: 38.4 / 76.4

    Qwen 3.5 122B Q3 K M: 32.9 / 68.6

    The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.

    I have done no further testing, as these results line up perfectly with my expectations.

    1. larodi · · focus · HN ↗
      To be done right, such vision benchmark should consider the peculiarities of the vision tower, and was it quantized and optimized. So a lot of success/loss of quality may not be due to the transformer.

      Second of all, the "find coordinates" of something is super difficult task of any model, so you tried to test a small quantized buddy with a tri-star challenge. Not sure what expectations were set.

      disclaimer: myself do large volume VLM work daily, including in production, for more than 1.5 years now.

      1. knollimar · · focus · HN ↗
        What type of work do you do? Do you consider shelling out to a tool where they can place markers and retry?

        Do you have any other benches? What models do you use or recommend for this type of work?

      2. Jackson__ · · focus · HN ↗
        >To be done right, such vision benchmark should consider the peculiarities of the vision tower

        As per my previous comment, I used the _exact same weights_ for the comparison. The vision tower is always kept as an unquantized BF16 file for GGUF, as I believe is the default.

        >Second of all, the "find coordinates" of something is super difficult task of any model

        It is in fact not (anymore), and most recent VLMs I tested have been standardized to point and bbox in a relative 0..1000 coordinate system with decent enough accuracy.

        And finally, none of this excuses that Strata performs so much worse with the same weights. Which does kinda leave me confused as to what the point of this reply is.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.