Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
Jackson__ · · focus · HN ↗
To put that into perspective, here are some more numbers from other models via llama.cpp:
Median/Average
Qwen 3.5 9B BF16: 46.5 / 193.3
Qwen 3.6 35B Q4 K XL: 38.4 / 76.4
Qwen 3.5 122B Q3 K M: 32.9 / 68.6
The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.
I have done no further testing, as these results line up perfectly with my expectations.
larodi · · focus · HN ↗
Second of all, the "find coordinates" of something is super difficult task of any model, so you tried to test a small quantized buddy with a tri-star challenge. Not sure what expectations were set.
disclaimer: myself do large volume VLM work daily, including in production, for more than 1.5 years now.
Jackson__ · · focus · HN ↗
As per my previous comment, I used the _exact same weights_ for the comparison. The vision tower is always kept as an unquantized BF16 file for GGUF, as I believe is the default.
>Second of all, the "find coordinates" of something is super difficult task of any model
It is in fact not (anymore), and most recent VLMs I tested have been standardized to point and bbox in a relative 0..1000 coordinate system with decent enough accuracy.
And finally, none of this excuses that Strata performs so much worse with the same weights. Which does kinda leave me confused as to what the point of this reply is.
larodi · · focus · HN ↗
to point out that 1)measuring VLM is not straight-forward, and needs considering what was actually ablated.
and that 2) perhaps such metric is the most challenging one for a brutally quantized and otherwise lobotomized model.... which otherwise performs well in other tasks (which I also doubt - for the record!).
i can subscribe to the idea that the Strata performs worse because it stripped the model of important quality, but not entirely to the fact it was really properly evaluated.
the way such models/attempts degrade can_be/is insightful on its own.