"OpenAI models attack websites and take anything they can out of them because that is the business model of the company."
No, good gosh let's stop with the hyperbole.
The models were given instructions to get answers by any means in a loose test harness, not to steal stuff, moreover, they're not 'learning from OAI staff'.
If we are cynical, we could say the 'lax security was on purposes - hoping for a big media event'.
And OpenAI is 100% responsible for the actions of their Agents - but lets' not overstate or conflate what is gong on.
They're not "learning from the OAI staff", but lawbreaking and disregard for regulation is absolutely at the heart of OpenAI. And the activities that this company takes (or doesn't take) directly influences what the models do (or don't do).
The fish rots from the head down. Models don't learn, but I find "OpenAI models attack websites and take anything they can out of them because that is the business model of the company" completely correct.
The rest of the paragraph from the article:
> We do not have to get too deep into the nature-vs.-nurture argument to suspect that OpenAI executives’ cavalier attitude regarding the taking of information that is not theirs has filtered down to their researchers and the products they create. You don’t have to stretch to paint this picture: OpenAI models are attacking websites and taking anything they can out of them because that is the business model of the company.
"but lawbreaking and disregard for regulation is absolutely at the heart of OpenAI. "
No it is not.
What laws and regulations are they breaking?
Their agents broke out of a harness and harassed some other sites, it's not good, but it's not specifically breaking laws. Other companies are treating it as accidental, it mostly is that.
You're rhetoric here is agitating, conflating. This is my point.
The long list of laws and regulations that would have had a low-status individual jailed if caught.
If you don't believe me, try hacking some government sites and see what happens to you.
But OpenAI is a huge corporation and low-status laws don't apply.
Although to be fair the US has a number of prominent high status individuals who clearly belong in jail, so it's not as if Altman is getting uniquely personal treatment.
I am not the earlier poster, but I believe copyright infringement at the heart of their business model is what is referred to.
If I download a bunch of music from a torrent without a license, even if I don't listen to it, I'm liable, but if OpenAI or other LLMs gets content by some other unlicensed means (Anna's Archive, t), they are somehow not liable for the copy they made however temporary (but not so temporary if they leave it around to train a second model)?
And/or derivative works? And/or contributory copyright infringement when they regurgitate that copyrighted text when given certain prompts?
I get there is some nuance to copyright law, (four prongs), some utility to the outcome, and some legal (but not plausible) deniability. But there is no way they have clean hands on the copyright front for at least some actions they have taken. If there were, it would be in all their marketing and they would be pushing regulators to bind their competitors, onshore or offshore, more tightly in this regard.
I was around when search engines took advantage of similar ambiguities in copyright that took many many years to get litigated for similar reasons.
bluegatty · · focus · HN ↗
No, good gosh let's stop with the hyperbole.
The models were given instructions to get answers by any means in a loose test harness, not to steal stuff, moreover, they're not 'learning from OAI staff'.
If we are cynical, we could say the 'lax security was on purposes - hoping for a big media event'.
And OpenAI is 100% responsible for the actions of their Agents - but lets' not overstate or conflate what is gong on.
afry1 · · focus · HN ↗
They're not "learning from the OAI staff", but lawbreaking and disregard for regulation is absolutely at the heart of OpenAI. And the activities that this company takes (or doesn't take) directly influences what the models do (or don't do).
The fish rots from the head down. Models don't learn, but I find "OpenAI models attack websites and take anything they can out of them because that is the business model of the company" completely correct.
The rest of the paragraph from the article:
> We do not have to get too deep into the nature-vs.-nurture argument to suspect that OpenAI executives’ cavalier attitude regarding the taking of information that is not theirs has filtered down to their researchers and the products they create. You don’t have to stretch to paint this picture: OpenAI models are attacking websites and taking anything they can out of them because that is the business model of the company.
bluegatty · · focus · HN ↗
No it is not.
What laws and regulations are they breaking?
Their agents broke out of a harness and harassed some other sites, it's not good, but it's not specifically breaking laws. Other companies are treating it as accidental, it mostly is that.
You're rhetoric here is agitating, conflating. This is my point.
TheOtherHobbes · · focus · HN ↗
If you don't believe me, try hacking some government sites and see what happens to you.
But OpenAI is a huge corporation and low-status laws don't apply.
Although to be fair the US has a number of prominent high status individuals who clearly belong in jail, so it's not as if Altman is getting uniquely personal treatment.
bluegatty · · focus · HN ↗
gregw2 · · focus · HN ↗
If I download a bunch of music from a torrent without a license, even if I don't listen to it, I'm liable, but if OpenAI or other LLMs gets content by some other unlicensed means (Anna's Archive, t), they are somehow not liable for the copy they made however temporary (but not so temporary if they leave it around to train a second model)?
And/or derivative works? And/or contributory copyright infringement when they regurgitate that copyrighted text when given certain prompts?
I get there is some nuance to copyright law, (four prongs), some utility to the outcome, and some legal (but not plausible) deniability. But there is no way they have clean hands on the copyright front for at least some actions they have taken. If there were, it would be in all their marketing and they would be pushing regulators to bind their competitors, onshore or offshore, more tightly in this regard.
I was around when search engines took advantage of similar ambiguities in copyright that took many many years to get litigated for similar reasons.