This is an excellent use case for completely local, small model inference, yet for inexplicable reasons Mozilla wants to normalize uploading your entire private browsing history to a cloud.
These (Mistral's and Mozilla's) marketing pages aren't candid enough to clearly explain the difference between local and cloud inference, and that they're asking you to consent to enabling the latter. I'd call that the bare minimum of ethics. If you're in a position of authority over your less technical users, you are ethically obligated to give them the full picture of what are you doing and why they should consent.
(Aside to any Mozilla people who might be reading HN, your page here[0] has an oversight, it advertises this model as "Mistral Small 4" but the hyperlink is to OpenAI's model card for gpt-oss-120b).
I am not trying to defend Mozilla doing this, and I don't support sending data to cloud based services like this in a way that users won't understand.
But I also think that the state of the art in small LLM and user device capabilities aren't there yet to put a "good enough to be actually useful" local-only LLM as a prepackaged thing in a mass market distributed browser.
You don't want a browser that takes 10GB of extra RAM (on top of the memory hog that is having just 3 or 4 complex tabs open on its own already) and pegs your CPU at 99% usage for minutes at a time. And not in an era when mass market consumer laptops are still commonly 8GB or 16GB of total system RAM. Many of those with integrated-into-CPU onboard graphics (eg: not a gaming laptop with a discrete GPU on the PCI-E bus).
It'll be a catastrophe for laptop battery life, among other resource use problems. And an LLM that fits in under 8 to 10GB of RAM for CPU-only inference is not going to be nearly as capable as an off-device inference system.
I wish they had just done this with a very clear up front opt in (not enabled by default) thing that explains what Mistral is, that it's not some big American cloud company but a relatively small startup in France, and that your prompts/LLM interactions will go to their servers. And some documentation on how it will be handled/stored in a supposedly trustworthy manner.
> that it's not some big American cloud company but a relatively small startup in France
How it's any better? Small companies can be bought by big companies. French government is as capable of trampling over their citizen's privacy the moment they feel they need it as US government is. And putting a squeeze on a small startup is way easier than on a major cloud company (not that either is particularly hard). Also, OpenAI used to be an idealistic non-profit one day too, then it started to smell trillions and all that went of of the window.
> is not going to be nearly as capable as an off-device inference system.
I rarely need PhD-level research into my browsing history. I'm not going to solve millennium problems on my bookmarks. The tasks that I will realistically need are well within capacity of most very basic local models. Maybe they'd be a bit slower, who cares.
Let's not pretend that EU data laws are not far more stringent and actually enforced than whatever the US has. Like sure, the government with a jury decision may acquire some person's data. But it's not getting sold to whichever-random-company-offered-the-highest.
Also, some shady TOS is also not enough to override the law in the EU, you have much more protection as a user here.
> But it's not getting sold to whichever-random-company-offered-the-highest.
Nothing in GDPR prevents transfer or selling of personal data. True, there are hoops to be jumped to do that, both in terms of consent and documentation, but let's not pretend a lot of people would read the welcome banner and refuse to interact with a service because its legalese says "we'll sell you data and if you don't want us to, go away". Yes, selling the data becomes more expensive because you need to hire the lawyers to produce those welcome pages and regulatory compliance paperwork, but large corps have enough money for lawyers.
> Also, some shady TOS is also not enough to override the law in the EU
It doesn't need to override anything, as the law does not prohibit data collection or transfer. It only describes the hoops that need to be jumped to achieve it.
It certainly does. You need to weigh the users privacy with your interests and can’t just put them aside. Noncompliance leads to quite large fines. You can easily find examples of fines given.
The DPA’s are short on capacity, sure, and large corps will try to fight the fines in court. But to describe it as “just hoops whilst allowing everything” is unfair.
> You need to weigh the users privacy with your interests and can’t just put them aside.
Who said anything about putting them aside? There would be a process and highly paid lawyers confirming that selling user data is only for the user's ultimate benefit, as it allows to provide awesome services to the users, and it all will be outlined in a 100-page privacy policy which you will be sure to read on each site you use, wouldn't you?
> But to describe it as “just hoops whilst allowing everything” is unfair.
Can you quote me the place where GDPR prohibits this? Not says something like "weigh user privacy" and "take adequate measures" and so on - which can always be resolved as "we weighed carefully and we took measures and we decided selling the data was awesome and users love it" - but explicitly and unambiguously prohibits the practice? If you don't find it - that description is exactly what it is.
peri-cl · · focus · HN ↗
These (Mistral's and Mozilla's) marketing pages aren't candid enough to clearly explain the difference between local and cloud inference, and that they're asking you to consent to enabling the latter. I'd call that the bare minimum of ethics. If you're in a position of authority over your less technical users, you are ethically obligated to give them the full picture of what are you doing and why they should consent.
(Aside to any Mozilla people who might be reading HN, your page here[0] has an oversight, it advertises this model as "Mistral Small 4" but the hyperlink is to OpenAI's model card for gpt-oss-120b).
[0] <a href="https://support.mozilla.org/en-US/kb/smart-window-models" rel="nofollow">https://support.mozilla.org/en-US/kb/smart-window-models
walrus01 · · focus · HN ↗
But I also think that the state of the art in small LLM and user device capabilities aren't there yet to put a "good enough to be actually useful" local-only LLM as a prepackaged thing in a mass market distributed browser.
You don't want a browser that takes 10GB of extra RAM (on top of the memory hog that is having just 3 or 4 complex tabs open on its own already) and pegs your CPU at 99% usage for minutes at a time. And not in an era when mass market consumer laptops are still commonly 8GB or 16GB of total system RAM. Many of those with integrated-into-CPU onboard graphics (eg: not a gaming laptop with a discrete GPU on the PCI-E bus).
It'll be a catastrophe for laptop battery life, among other resource use problems. And an LLM that fits in under 8 to 10GB of RAM for CPU-only inference is not going to be nearly as capable as an off-device inference system.
I wish they had just done this with a very clear up front opt in (not enabled by default) thing that explains what Mistral is, that it's not some big American cloud company but a relatively small startup in France, and that your prompts/LLM interactions will go to their servers. And some documentation on how it will be handled/stored in a supposedly trustworthy manner.
smsm42 · · focus · HN ↗
How it's any better? Small companies can be bought by big companies. French government is as capable of trampling over their citizen's privacy the moment they feel they need it as US government is. And putting a squeeze on a small startup is way easier than on a major cloud company (not that either is particularly hard). Also, OpenAI used to be an idealistic non-profit one day too, then it started to smell trillions and all that went of of the window.
> is not going to be nearly as capable as an off-device inference system.
I rarely need PhD-level research into my browsing history. I'm not going to solve millennium problems on my bookmarks. The tasks that I will realistically need are well within capacity of most very basic local models. Maybe they'd be a bit slower, who cares.
gf000 · · focus · HN ↗
Also, some shady TOS is also not enough to override the law in the EU, you have much more protection as a user here.
smsm42 · · focus · HN ↗
Nothing in GDPR prevents transfer or selling of personal data. True, there are hoops to be jumped to do that, both in terms of consent and documentation, but let's not pretend a lot of people would read the welcome banner and refuse to interact with a service because its legalese says "we'll sell you data and if you don't want us to, go away". Yes, selling the data becomes more expensive because you need to hire the lawyers to produce those welcome pages and regulatory compliance paperwork, but large corps have enough money for lawyers.
> Also, some shady TOS is also not enough to override the law in the EU
It doesn't need to override anything, as the law does not prohibit data collection or transfer. It only describes the hoops that need to be jumped to achieve it.
tinodb · · focus · HN ↗
The DPA’s are short on capacity, sure, and large corps will try to fight the fines in court. But to describe it as “just hoops whilst allowing everything” is unfair.
smsm42 · · focus · HN ↗
Who said anything about putting them aside? There would be a process and highly paid lawyers confirming that selling user data is only for the user's ultimate benefit, as it allows to provide awesome services to the users, and it all will be outlined in a 100-page privacy policy which you will be sure to read on each site you use, wouldn't you?
> But to describe it as “just hoops whilst allowing everything” is unfair.
Can you quote me the place where GDPR prohibits this? Not says something like "weigh user privacy" and "take adequate measures" and so on - which can always be resolved as "we weighed carefully and we took measures and we decided selling the data was awesome and users love it" - but explicitly and unambiguously prohibits the practice? If you don't find it - that description is exactly what it is.