Qwen 3.8 Omni Flash
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Qwen 3.8 Omni Flash
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
tolugenius · · focus · HN ↗
_ache_ · · focus · HN ↗
The last one was: Qwen3-Omni-30B-A3B <a href="https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct" rel="nofollow">https://huggingface.co/Qwen/Qwen3-Omni-30B-A3B-Instruct
And maybe Qwen4 won't be released, they only release Qwen3.8 27B (and a mostly unusable 125B). There are definitively slowing down open weight release.
imrehg · · focus · HN ↗
bitexploder · · focus · HN ↗
nicman23 · · focus · HN ↗
[dead]
_ache_ · · focus · HN ↗
diddid · · focus · HN ↗
mdp2021 · · focus · HN ↗
Surely it was meant to be 'definitely' - the "good news" at this stage are that given the speed of history and important levels of uncertainty, it is difficult to label trends with "definitively" ;)
Iolaum · · focus · HN ↗
I assume you are talking about qwen3.8-flash-next. Support for it on some places, like llama.cpp, is still wip (depending on configuration) but it looks like a very capable model in it's category.
_ache_ · · focus · HN ↗
So, the cost of a setup to run Qwen-Flash-Next at +40tks is around $3000. Too much for most people.
With only a RTX 4090, you will reach 30tps (with DDR5...), not +40tks, and it's about the limit to be usable. Oh ! I forget Apple device too, it's a good option to run this model I guess, but still slow.
Yet, as you said, it's still a wip implementation, it may improve soon (MTP support is about to be merged in llama.cpp soon).
vinzenzu · · focus · HN ↗
I still agree that they aren't as aggressively releasing the open-weights models as before, but there hasn't been a major release they haven't published the weights for yet afaik.
[1] <a href="https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B" rel="nofollow">https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B [2] <a href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next" rel="nofollow">https://huggingface.co/Qwen/Qwen3.8-Flash-Next
npodbielski · · focus · HN ↗
hgoel · · focus · HN ↗
conception · · focus · HN ↗
rubslopes · · focus · HN ↗
What would that mean in this context?
cleaning · · focus · HN ↗
smallerfish · · focus · HN ↗
pennomi · · focus · HN ↗
I swear I spend more time telling Claude not to do things than telling it what to do.
mdp2021 · · focus · HN ↗
But is that because of training, or can that be (also? mostly?) an effect of the system prompt?
girvo · · focus · HN ↗
vintermann · · focus · HN ↗
disgruntledphd2 · · focus · HN ↗
Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.
khafra · · focus · HN ↗
Reinforcement Learning (in LLMs) trains via gradient descent on a reward signal that's an imperfect proxy for the actual goal of the engineers doing the training. So, under mild optimization pressure, you get increasingly more of what you want, because that's the easiest way to increase the metric.
But as the optimization pressure increases, so do the ways to increase the metric by doing increasingly weird things. If the full action space grows sufficiently faster than the "things you actually want" subset, the amount of "things you actually want" goes to 0 under sufficient RL.
conception · · focus · HN ↗
antupis · · focus · HN ↗
spijdar · · focus · HN ↗
I've noticed between tool calls, it'll sometimes say things like:
These don't clearly reflect ... anything, and it keeps performing tool calls correctly anyway. And then other times, it begins doing whatever you'd call this (this is only orthogonally related to the task):nojs · · focus · HN ↗
Morizero · · focus · HN ↗
denom · · focus · HN ↗
Wow, that is unexpected. But honest?
nine_k · · focus · HN ↗
Bluestein · · focus · HN ↗
kouteiheika · · focus · HN ↗
Is this with the full unquantized weights? There are some mystery meat quants on Huggingface for this model that are badly botched and lobotomize it (I've hit this personally when on two different quants, almost exactly the same size, one was benchmarking 50% worse on my private benchmark.).
spijdar · · focus · HN ↗
seemaze · · focus · HN ↗
ryan-c · · focus · HN ↗
saghm · · focus · HN ↗
Bluestein · · focus · HN ↗
saghm · · focus · HN ↗
Bluestein · · focus · HN ↗
saghm · · focus · HN ↗
> Error: Thank you for participating in the Stealth Union Alpha testing period. This model was Unbiased's Pareto.
One of my friends quipped that "Unbiased Pareto" still sounded like the name of a stealth model.
I'm not sure if I just noticed it later, but this definitely seemed to be a lot shorter than other stealth alphas I've tried. I wouldn't be shocked if this is more typical going forward though, or if stealth alphas entirely go away, since it certainly costs a bit of money to market this way.
Bluestein · · focus · HN ↗
We got drive-by modelled.-
saghm · · focus · HN ↗
I had the $20/month Claude one for a few months starting in February, but one day it randomly started returning me errors claiming I needed to pay for more credits despite the usage showing 8% for the week and 20% for the session, and I figured if they couldn't even communicate to me the difference between them screwing up the check for hitting the limit or an outage, it wasn't worth it for me to keep paying them. I dislike OpenAI too much to want to pay them any of my personal money for anything, and when I tried out Mistral Vibe it did not work very well for me (it kept not following instructions and eventually when I kept trying to push it to handle things better it somehow spiraled into simulating some sort of existential crisis, culminating in gibberish and random characters being dumped on my screen infinitely until I killed the process; incredibly entertaining, but not worth paying for)
Bluestein · · focus · HN ↗
Sensible.-
pigeons · · focus · HN ↗
zozbot234 · · focus · HN ↗
anon373839 · · focus · HN ↗
Another failure mode you may see is inordinately long CoT. Properly served, the model is good at calibrating its CoT length to the difficulty of the immediate task.
[1] <a href="https://github.com/blazux/qwen3.8-Flash-DGX" rel="nofollow">https://github.com/blazux/qwen3.8-Flash-DGX
ryan-c · · focus · HN ↗
hgoel · · focus · HN ↗
girvo · · focus · HN ↗
Still works great though!
whizzter · · focus · HN ↗
syntaxing · · focus · HN ↗
Wow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of people. Other capability also matches or exceeds 3.8 Flash.
They also made a new harness but github link seems to 404.
testaburger · · focus · HN ↗
JLO64 · · focus · HN ↗
[dead]
_ache_ · · focus · HN ↗
in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47
That is a massive cost reduction.
Refs: <a href="https://www.alibabacloud.com/help/en/model-studio/model-pricing#china-beijing-h4" rel="nofollow">https://www.alibabacloud.com/help/en/model-studio/model-pric... <a href="https://runware.ai/gemini-omni" rel="nofollow">https://runware.ai/gemini-omni
killingtime74 · · focus · HN ↗
vntok · · focus · HN ↗
kaliqt · · focus · HN ↗
LeBit · · focus · HN ↗
Even if OpenAI end up using 1 token for per task, if the token costs 1M$ , some people will find it expensive.
nater5000 · · focus · HN ↗
Astra on xhigh has a cost per task of $2.31 with an intelligence index of 53. Qwen3.8 Max has a cost per task of $5.41 with an intelligence index of 45. Pricing for GPT-6 Astra (xhigh) is $10.00 per 1M input tokens and $50.00 per 1M output tokens. Pricing for Qwen3.8 Max (0902) is $2.00 per 1M input tokens and $6.00 per 1M output tokens.
Obviously this is just one measure of all of this (and Qwen 3.8 Omni Flash isn't yet available), but I think this illustrates the point well. These relative task costs are pretty consistent across different analysts. Cost per token is arguably a useless measure at this point in most circumstances.
alphabettsy · · focus · HN ↗
tidbeck · · focus · HN ↗
npodbielski · · focus · HN ↗
_ache_ · · focus · HN ↗
nater5000 · · focus · HN ↗
If Gemini can complete a task for $1 and Qwen completes that same task for $1, then the cost per token is irrelevant in most use-cases. One would think this stuff should correlate well enough that you can use it as a proxy, but I think a lot of people are noticing this is a serious mistake and that these "cheap" models aren't as cheap as they appear when you consider this.
_ache_ · · focus · HN ↗
What matter the most and isn't told by token price is the latency. You expect a voice LLM to respond very quick. If it takes 5s to response to a simple "Hello, what the weather today?", them not much people will use it.
lxe · · focus · HN ↗
andy_ppp · · focus · HN ↗
The Chinese just seem to have an ability to get it done without anywhere near the GPUs of the US and Europe can buy these GPUs.
I think relying on the US and China for AI is probably not ideal? For example I think Qwen have not released the Omni models as open weights in the past, it’d be good to know if they’re doing this here?
apexalpha · · focus · HN ↗
We have no such government with a mandate to do that.
EagnaIonat · · focus · HN ↗
The main one (IIRC) is Spains ALIA.
<a href="https://alia.gob.es/eng" rel="nofollow">https://alia.gob.es/eng
There is also OpenEuroLLM and EuroLLM.
<a href="https://www.openeurollm.eu" rel="nofollow">https://www.openeurollm.eu
jambutters · · focus · HN ↗
Zambyte · · focus · HN ↗
podocarp · · focus · HN ↗
rw2 · · focus · HN ↗
Europe has none;
The best tech university in Europe when compared to Chinese/US equivalents won't even rank in the top 10.
bitexploder · · focus · HN ↗
EagnaIonat · · focus · HN ↗
I feel people are too focused on US/China and don't pay attention to what is going on in the world.
ASML in the Netherlands for example was the only company in the world that makes EUV lithography machines, which all the major chip companies depend on. China recently reverse engineered their work to create machines since late 2025, but not sold commercially.
Ireland has chip production facilities.
EU might not be in the top 2, but it is not out of the running at all.
bitexploder · · focus · HN ↗
hshdjdjdif · · focus · HN ↗
trvz · · focus · HN ↗
The LLM they were involved in last year (Apertus) was a letdown.
As someone who actually went there, my impression is that Europe in general is complacent when it comes to computers, and any bright eyed student will get their motivation choked out of them in academia here.
If you want to do things with AI in Europe, you can have a bigger effect by working for a consulting company than being at a university. That's … not a good sitatution.
danielscrubs · · focus · HN ↗
Europeans do not have the hustle mentality to break the law like Uber so you wont see them compete in anything data heavy. They also are quite risk adverse, probably from having a lower Gini coefficient.
le-mark · · focus · HN ↗
This plus the lower Gini coefficient is more a comment on taxation and capital investment. Why bother as founders when early round investors and taxes will take it all?
AIiscoming · · focus · HN ↗
EagnaIonat · · focus · HN ↗
Well there is Mistral. The EU is ahead on specialised models than general purpose LLMs.
There is.
- Flux3
- Kyutai (Open source AI lab)
- H Company
- LightOn
- AMD Silo (Finland)
- OpenEuroLLM and EuroLLM
There is probably more, but that's off the top of my head.
miohtama · · focus · HN ↗
EagnaIonat · · focus · HN ↗
Black Forest Labs for example has a $4B valuation with half a billion raised so far.
miohtama · · focus · HN ↗
[dead]
ffsm8 · · focus · HN ↗
Europe actually follows americas rules, hence they're not doing that.
It's braindead for sure considering how the US treats europe, but it is what it's
miohtama · · focus · HN ↗
Because any AI company would be hit by hate and regulation derived from this, no VC invests in the EU. It's more state and large enterprise investments, good old East Germany style. And historically it has not been that efficient.
Or simply: lack of risk taking appetite.
AIiscoming · · focus · HN ↗
No?
Its just that the richest companys with the most VC sit in USA and Europe isn't used to pay what USA / VC is paying and we are a little bit slow.
andy_ppp · · focus · HN ↗
miohtama · · focus · HN ↗
zipy124 · · focus · HN ↗
nananana9 · · focus · HN ↗
How many times will we do this? This line of thought is precisely why we can't build almost anything, and China can built almost everything. We could do X, but we're too smart - let's offload the actual work to China and we'll run our economy on IP and B2B deals.
esperent · · focus · HN ↗
nullbio · · focus · HN ↗
nullbio · · focus · HN ↗
podocarp · · focus · HN ↗
spacebanana7 · · focus · HN ↗
drbscl · · focus · HN ↗
xutopia · · focus · HN ↗
Omni means you can use multiple types of input and have multiple types of outputs like audio, video, images and text. Flash means that it is built for speed and smaller than the more complete ones.
mavamaarten · · focus · HN ↗
I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.
E.g. I wrote a tool that cleans out my email spam box. It classifies emails that are already flagged as spam, and if it's very obviously spam it removes it permanently (keeps a copy on disk though). And after x emails, it goes through the list of deleted spam mails and suggests email rules. What model would be best suited? I'd love to be able to explain this use case and get this info served to me. The list of models and the information about what they're good at is just too splintered and spread out. I landed on google/gemma-4-31b for now, because it's cheap and good enough and also supports Dutch and French a bit. But I can't realistically try them all.
apimade · · focus · HN ↗
But this is only stuff I can run locally, or it’s a cloud model.
So.. This is good for right now!
<a href="https://apimade.com/audio-compare.html" rel="nofollow">https://apimade.com/audio-compare.html
Systemerror7A69 · · focus · HN ↗
As long as the model you're using solves the problems you have to your satisfaction, there is no need to try any other models, except for financial reasons maybe.
So I start with a relatively cheap model (GLM 5.3 flash for me) and as long as it accomplishes the task (it did so far) I don't have to change. And even if it can't do something, the first thing I change is see if I can give it more tools or better context (useful even if I switch models later) or trying a different approach to the problem.
If google/gemma-4-31b works, you don't need to overthink it.
blensor · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
insightfulornot · · focus · HN ↗
alfiedotwtf · · focus · HN ↗
epolanski · · focus · HN ↗
Up until recently I had a gemini flash 2.0 api deployed that did summarization and translation of news articles/corporate statements fast and cheap and had no reason to update it.
If it works fine, this chase of the latest LLM is bit pointless.
Lalabadie · · focus · HN ↗
(I think you're right)
verst · · focus · HN ↗
cogman10 · · focus · HN ↗
It also helps that a lot of effort has already been spent figuring out how to do more with weaker models because SOTA 1 year ago was behind what the cheap models do today.
1 year from now, unless the SOTA companies come up with something truly revolutionary they will be in a lot of trouble.
t_mahmood · · focus · HN ↗
Yeah, might be far fetched, but it seems like heading that way
bmordue · · focus · HN ↗
barrenko · · focus · HN ↗
Long term I have fears I can't depend of the Google's AI api.
killingtime74 · · focus · HN ↗
sisve · · focus · HN ↗
<a href="https://openrouter.ai/docs/cookbook/coding-agents/openclaw-integration#using-auto-model-for-cost-optimization" rel="nofollow">https://openrouter.ai/docs/cookbook/coding-agents/openclaw-i...
ComputerGuru · · focus · HN ↗
mavamaarten · · focus · HN ↗
croon · · focus · HN ↗
Mstly just popularity trends, hoping that there's some wisdom in the crowd.
deedree · · focus · HN ↗
bonoboTP · · focus · HN ↗
They generally all try to make them good at everything, it's not like they'd declare "this model is not made for task X".
edude03 · · focus · HN ↗
ledak · · focus · HN ↗
Ask one of the top-tier models to do deep research on it.
rudicjd2747 · · focus · HN ↗
yorwba · · focus · HN ↗
Additionally, using a Chinese provider doesn't mean the model will be hosted in China, e.g. for Qwen Omni here, the supported regions are: China (Beijing), Singapore, China (Hong Kong), Japan (Tokyo), Germany (Frankfurt), and US (Virginia). <a href="https://www.alibabacloud.com/help/en/model-studio/qwen-omni#qwen3-8-omni-flash" rel="nofollow">https://www.alibabacloud.com/help/en/model-studio/qwen-omni#...
boltzmann-brain · · focus · HN ↗
that aside, solar also includes storage of energy that comes from solar. you should read up on energy networks.
idiotsecant · · focus · HN ↗
paimapi · · focus · HN ↗
that said, China is also rapidly scaling up coal-fired plants: <a href="https://apnews.com/article/china-coal-power-plant-carbon-climate-change-ba86e7584e3afe1826eed5cffa25354a" rel="nofollow">https://apnews.com/article/china-coal-power-plant-carbon-cli...
those presumably support all of the surrounding infrastructure + people + manufacturing so it's not as if it's truly solar-powered. but it's still handily better than the state-by-state abandonment of clean energy goals here in the US - I lay this out a bit here: <a href="https://news.ycombinator.com/item?id=49700743">https://news.ycombinator.com/item?id=49700743
yorwba · · focus · HN ↗
paimapi · · focus · HN ↗
this is a pedantic misinterpretation of the scope of East Data, West Computing. the hubs are all in tier one cities with the power generation capacity to handle them. each hub is comprised of dozens of data-centers and are the place to build because of large economic incentives, infra guarantees, and university-trained talent living in close proximity
see <a href="https://www.sciencedirect.com/science/article/pii/S2095809924005058" rel="nofollow">https://www.sciencedirect.com/science/article/pii/S209580992...
>This initiative is expected to accommodate up to 95% of China’s digital data needs through this reorganization
>Within a year of implementation, more than 112 new data centers have either been authorized, under construction, or completed across the planned computing hubs
this project itself is the thing that has all the other nation states in the world clutching pearls about not supporting their native AI corporations enough. it's something you and everyone else should really familiarize yourself with because it is one of the major items influencing current realpolitik and economic decisions (and is likely going to be the thing that'll lead us to a thrice-in-a-lifetime sized recession but that's a convo for a different day)
ezst · · focus · HN ↗
If by that you misspelled coal, sure
<a href="https://ourworldindata.org/grapher/share-elec-by-source?country=~CHN" rel="nofollow">https://ourworldindata.org/grapher/share-elec-by-source?coun...
defrost · · focus · HN ↗
ezst · · focus · HN ↗
defrost · · focus · HN ↗
Kind of like what absolutely isn't happening with the current US administration.
Fizz43 · · focus · HN ↗
defrost · · focus · HN ↗
Check the source we're discussing: <a href="https://ourworldindata.org/grapher/share-elec-by-source?country=~CHN" rel="nofollow">https://ourworldindata.org/grapher/share-elec-by-source?coun...
As you can see it is percentage of coal as part of total energy mix that has been dropping for two decades.
As already stated above. Perhaps you missed that?
You're correct that absolute tonnage of coal use is still climbing, - but at an ever decreasing rate and was predicted to peak .. of course, thanks to some shenanigans by the Hegseth's of the world that may be deferred for a while, thanks USofA!
Of course on the plus side (for the atmosphere) there's been a drop in global fossil fuel consumption, so swings, arrows.
China, of course, is still well short of total CO2 tonnage lofted over the past century in comparison with the US - even allowing for a vastly greater population.
Fizz43 · · focus · HN ↗
testycool · · focus · HN ↗
<sidenote>
Similarly, HuggingFace has a CLI + a few skills, and they are very useful.
I had a production image processing using Gemini 2.5 Flash Lite (which is getting discontinued in October), and in 20 minutes Claude Code + HF Cli recommended the best replacement small model (Qwen VL 3B something) and proceeded to fine tune it on my datataset. All this while I was in a rush to get dressed and go to the store.
It cost ~$3 I think, and results were excellent. Not perfect, but not far from perfect either.
We didn't replace Gemini in prod at the time, because we didn't have time to do all the math on how to end up with a smaller bill/mo.
</sidenote>
apples_oranges · · focus · HN ↗
stevenhubertron · · focus · HN ↗
greazy · · focus · HN ↗
bushido · · focus · HN ↗
And that has any models that I'm considering through that.
maxyurk · · focus · HN ↗
revenue_sub_doe · · focus · HN ↗
[dead]
revenue_sub_doe · · focus · HN ↗
[dead]
revenue_sub_doe · · focus · HN ↗
[dead]
esquire_900 · · focus · HN ↗
zepearl · · focus · HN ↗