This is an excellent use case for completely local, small model inference, yet for inexplicable reasons Mozilla wants to normalize uploading your entire private browsing history to a cloud.
These (Mistral's and Mozilla's) marketing pages aren't candid enough to clearly explain the difference between local and cloud inference, and that they're asking you to consent to enabling the latter. I'd call that the bare minimum of ethics. If you're in a position of authority over your less technical users, you are ethically obligated to give them the full picture of what are you doing and why they should consent.
(Aside to any Mozilla people who might be reading HN, your page here[0] has an oversight, it advertises this model as "Mistral Small 4" but the hyperlink is to OpenAI's model card for gpt-oss-120b).
I am not trying to defend Mozilla doing this, and I don't support sending data to cloud based services like this in a way that users won't understand.
But I also think that the state of the art in small LLM and user device capabilities aren't there yet to put a "good enough to be actually useful" local-only LLM as a prepackaged thing in a mass market distributed browser.
You don't want a browser that takes 10GB of extra RAM (on top of the memory hog that is having just 3 or 4 complex tabs open on its own already) and pegs your CPU at 99% usage for minutes at a time. And not in an era when mass market consumer laptops are still commonly 8GB or 16GB of total system RAM. Many of those with integrated-into-CPU onboard graphics (eg: not a gaming laptop with a discrete GPU on the PCI-E bus).
It'll be a catastrophe for laptop battery life, among other resource use problems. And an LLM that fits in under 8 to 10GB of RAM for CPU-only inference is not going to be nearly as capable as an off-device inference system.
I wish they had just done this with a very clear up front opt in (not enabled by default) thing that explains what Mistral is, that it's not some big American cloud company but a relatively small startup in France, and that your prompts/LLM interactions will go to their servers. And some documentation on how it will be handled/stored in a supposedly trustworthy manner.
I'm worried that the middle could fall out of the computing market across the board. If you can afford to keep up with the upgrade treadmill, you'll get private, local inference capabilities. If you can't afford to stay on the treadmill, you'll be stuck with whatever cloudshit malware Silicon Valley wants to foist on you.
I acknowledge that this is already the case, to some extent. The cheapest laptops at Best Buy are crammed with the most preinstalled malware. That's been the case for, what, 25 years? But you've always been able to wipe that cheap laptop and make it into a much more capable, trustworthy machine.
Well, assuming LLMs do become a pervasive part of the computing experience, what happens to the cheap laptops? Do all computers get more expensive to accommodate local inference? Does the rift between the everyday user's experience and the savvy user's experience grow even wider than it already is? Neither outcome seems good for the average joe who just needs to check his email.
Remember how in like 1999/2000 Sun was trying to predict that everyone's computer would be some form of thin terminal in the future? Turns out they were very wrong on the part about it running on Sun server back-end infrastructure, but that same general purpose has now been accomplished through other methods where a lot of people do basically EVERYTHING inside a web browser tab to some external cloud service.
Now add the need for external inference because very few random consumers are going to buy a $3000 laptop when they can get the $600 laptop at Best Buy, and that trend further escalates.
This has always been the case, forever. You have to pay for a product or service. How you do so can be with cash or your data/body/vote/eyeballs/indirect discretionary purchases.
The amount of work that can be done funded by foundations and free work is nowhere close to what people want.
Oh give me a break. This argument that people's objections to advertising comes from some Pollyanna naivety over things being free is nonsense. Tell me the last time you have ever seen a company be upfront and explicitly offer a free tier where they tell you upfront exactly what they are harvesting about you and for what purpose (and no, "improving user experience" isn't being upfront) but also offer you a paid version where they explicitly promise not to do that.
"If the product is free then you are the product" is supposed to be a cautionary observation, not an axiomatic proscription.
Does anyone actually trust companies that offer paid services to not harvest their data? Whenever I see this trope I think to myself that whatever lack of regulation, oversight and enforcement lead to that being okay would equally allow for them to both take my money and harvest all my of data anyways.
This is akin to people who criticise socialised medicine or services and pejoratively characterize supporters as just wanting "free stuff", as if the concept of collective payment is naive or something
peri-cl · · focus · HN ↗
These (Mistral's and Mozilla's) marketing pages aren't candid enough to clearly explain the difference between local and cloud inference, and that they're asking you to consent to enabling the latter. I'd call that the bare minimum of ethics. If you're in a position of authority over your less technical users, you are ethically obligated to give them the full picture of what are you doing and why they should consent.
(Aside to any Mozilla people who might be reading HN, your page here[0] has an oversight, it advertises this model as "Mistral Small 4" but the hyperlink is to OpenAI's model card for gpt-oss-120b).
[0] <a href="https://support.mozilla.org/en-US/kb/smart-window-models" rel="nofollow">https://support.mozilla.org/en-US/kb/smart-window-models
walrus01 · · focus · HN ↗
But I also think that the state of the art in small LLM and user device capabilities aren't there yet to put a "good enough to be actually useful" local-only LLM as a prepackaged thing in a mass market distributed browser.
You don't want a browser that takes 10GB of extra RAM (on top of the memory hog that is having just 3 or 4 complex tabs open on its own already) and pegs your CPU at 99% usage for minutes at a time. And not in an era when mass market consumer laptops are still commonly 8GB or 16GB of total system RAM. Many of those with integrated-into-CPU onboard graphics (eg: not a gaming laptop with a discrete GPU on the PCI-E bus).
It'll be a catastrophe for laptop battery life, among other resource use problems. And an LLM that fits in under 8 to 10GB of RAM for CPU-only inference is not going to be nearly as capable as an off-device inference system.
I wish they had just done this with a very clear up front opt in (not enabled by default) thing that explains what Mistral is, that it's not some big American cloud company but a relatively small startup in France, and that your prompts/LLM interactions will go to their servers. And some documentation on how it will be handled/stored in a supposedly trustworthy manner.
ryukoposting · · focus · HN ↗
I acknowledge that this is already the case, to some extent. The cheapest laptops at Best Buy are crammed with the most preinstalled malware. That's been the case for, what, 25 years? But you've always been able to wipe that cheap laptop and make it into a much more capable, trustworthy machine.
Well, assuming LLMs do become a pervasive part of the computing experience, what happens to the cheap laptops? Do all computers get more expensive to accommodate local inference? Does the rift between the everyday user's experience and the savvy user's experience grow even wider than it already is? Neither outcome seems good for the average joe who just needs to check his email.
walrus01 · · focus · HN ↗
Now add the need for external inference because very few random consumers are going to buy a $3000 laptop when they can get the $600 laptop at Best Buy, and that trend further escalates.
tyre · · focus · HN ↗
The amount of work that can be done funded by foundations and free work is nowhere close to what people want.
gremlinunderway · · focus · HN ↗
"If the product is free then you are the product" is supposed to be a cautionary observation, not an axiomatic proscription.
Does anyone actually trust companies that offer paid services to not harvest their data? Whenever I see this trope I think to myself that whatever lack of regulation, oversight and enforcement lead to that being okay would equally allow for them to both take my money and harvest all my of data anyways.
This is akin to people who criticise socialised medicine or services and pejoratively characterize supporters as just wanting "free stuff", as if the concept of collective payment is naive or something