Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
Jackson__ · · focus · HN ↗
To put that into perspective, here are some more numbers from other models via llama.cpp:
Median/Average
Qwen 3.5 9B BF16: 46.5 / 193.3
Qwen 3.6 35B Q4 K XL: 38.4 / 76.4
Qwen 3.5 122B Q3 K M: 32.9 / 68.6
The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.
I have done no further testing, as these results line up perfectly with my expectations.
NamlchakKhandro · · focus · HN ↗
unlikelytomato · · focus · HN ↗
Borealid · · focus · HN ↗
It's just easier to measure the "right" answer (and deviation therefrom) on a vision task than a language one due to the underspecified nature of language.
unlikelytomato · · focus · HN ↗
dr_kiszonka · · focus · HN ↗
pjc50 · · focus · HN ↗
throwaway219450 · · focus · HN ↗
Xenograph · · focus · HN ↗
biztos · · focus · HN ↗
larodi · · focus · HN ↗
Second of all, the "find coordinates" of something is super difficult task of any model, so you tried to test a small quantized buddy with a tri-star challenge. Not sure what expectations were set.
disclaimer: myself do large volume VLM work daily, including in production, for more than 1.5 years now.
knollimar · · focus · HN ↗
Do you have any other benches? What models do you use or recommend for this type of work?