Frontier Labs Are Selling Garbage to Fools in Washington
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Frontier Labs Are Selling Garbage to Fools in Washington
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
anigbrowl · · focus · HN ↗
jgalt212 · · focus · HN ↗
every layperson is an expert in 1 or more areas.
randallsquared · · focus · HN ↗
This isn't even remotely true, unless you're willing to go as far as "being this person is an area of expertise".
sublinear · · focus · HN ↗
I suppose you've never had a unique perspective on something? How unfortunate.
basch · · focus · HN ↗
Now if they taught them to express uncertainty and speak in terms of probability, I might be more likely to be blindly fooled by some kind of uncertain conviction.
jgalt212 · · focus · HN ↗
That's quite an extraordinary claim there which I doubt would hold up to an even cursory level of scrutiny. But if you yell it loud enough, maybe people will be afraid to challenge your assertion.
iLoveOncall · · focus · HN ↗
What do you think 70 years old career politicians are an expert in that would allow them to weigh whether a chatbot's answers are bullshit or not.
sublinear · · focus · HN ↗
That's their expertise.
tracerbulletx · · focus · HN ↗
AngryData · · focus · HN ↗
Our average politician is far above the age that scammers target with LLM based scam phone calls, and being a politician with wealth insulates them from reality and most consequences already, they are completely out of their depth from every direction.
hackernews682 · · focus · HN ↗
lokar · · focus · HN ↗
DalasNoin · · focus · HN ↗
This is incorrect, the HF incident for example (the most well known) had nothing to do with irregular. I know there has been a news site pushing inaccurate articles (effort.news) on this topic but these are the facts.
<a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/" rel="nofollow">https://openai.com/index/hugging-face-incident-and-the-road-...
nr378 · · focus · HN ↗
kalkin · · focus · HN ↗
> For Anthropic, Google, and Meta, the catastrophic breakouts happened inside the testing environments of the exact same contractor.
If this is the level of understanding you have of the relevant incidents, there's a lot of chutzpah in saying that other people are "selling garbage", carrying out an "extraordinary confidence trick", etc.
nr378 · · focus · HN ↗
[dead]
DalasNoin · · focus · HN ↗
aesthesia · · focus · HN ↗
verdverm · · focus · HN ↗
genuinely curious, haven't heard others raise any yet, but does not mean it is issue free
DalasNoin · · focus · HN ↗
bigstrat2003 · · focus · HN ↗
0xDEAFBEAD · · focus · HN ↗
<a href="https://finance.yahoo.com/technology/ai/articles/anthropic-tops-100-billion-revenue-224001996.html" rel="nofollow">https://finance.yahoo.com/technology/ai/articles/anthropic-t...
I don't think people pay that sort of money for technology which is useless.
AlexCoventry · · focus · HN ↗
[1] <a href="https://openai.com/index/navier-stokes-solution/" rel="nofollow">https://openai.com/index/navier-stokes-solution/
etcetcetcetceta · · focus · HN ↗
Do you feel like you're getting good value for money if you solve a millenium problem at a fourteen million dollar loss?
another_twist · · focus · HN ↗
Yes you do. If we factor in wages assigned to skilled mathematicians who have worked on this problem, then the 14m for a 1mUSD prize is good value. Also, once things are solved, it on to applications where even more value could potentially be generated. Here its 1m USD simply because theres a proze assigned to it. But once these solution are used in applications, the yield is calculated on the 14m RnD cost. Which (depending on the application) is a good investment.
pliny · · focus · HN ↗
nr378 · · focus · HN ↗
pliny · · focus · HN ↗
[1] <a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf" rel="nofollow">https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78... - page 9
nr378 · · focus · HN ↗
Chaining a public token to an 11-year-old Jinja2 template injection vuln shouldn't be dressed up as an unprecedented "alien intellect" that threatens human civilisation. (And HuggingFace should take some flack for having such a dated vulnerability exposed - if your Bank was compromised in this way, you'd be blaming your bank, not the attacker.)
One correction is fair though, the 14 tokens were in a public Hugging Face dataset not a public GitHub repository. I've updated the post to reflect that.
[1] <a href="https://blackhat.com/docs/us-15/materials/us-15-Kettle-Server-Side-Template-Injection-RCE-For-The-Modern-Web-App-wp.pdf" rel="nofollow">https://blackhat.com/docs/us-15/materials/us-15-Kettle-Serve...
franga2000 · · focus · HN ↗
> This is one of the most lucid pieces of writing capturing the current state of play I’ve read. Who is the author?
verdverm · · focus · HN ↗
tyranny of confirmational headlines
tim333 · · focus · HN ↗
EA-3167 · · focus · HN ↗
<a href="https://www.pangram.com/history/6451ec6b-90b6-4e17-bfc9-6730011b1af4?ucc=I4Fo2JL9GdU" rel="nofollow">https://www.pangram.com/history/6451ec6b-90b6-4e17-bfc9-6730...
skeledrew · · focus · HN ↗
trhway · · focus · HN ↗
Currently we have a classic market race - the market forces both companies to provide more and more capable models with more and more value for the money for the customers while the suppliers (Nvidia, Micron, Dell) squeeze them from the other side. The end result would be one winner taking it all (and that is a very humongous "all") while the other falling into a very distant second position at best. Do their leadership and investors (bonus points - look at some OpenAI investors), some already worth tens of billions on paper, like the prospects of that "50% chances of even more riches, 50% - bust" outcome? I'd think - no.
The outcome they would like is both companies divvying up that huge market while jointly raising prices and providing less capable models (ie. cheaper to train and run) while squeezing their suppliers Wallmart style. How to get there? By breaking the anti-cartel limitations.
The typical tools to break anti-cartel limitations is for example perception of national interests or perception of some imminent dire emergency.
Thus the "lets us collaborate or our AI will kill you all" scare campaign.
defgeneric · · focus · HN ↗
They've been raising the issue bi-monthly through mini-scandals that have until now been consistently slapped down by Jensen Huang, but it seems an alliance with Democrats, just prior to an election where they're poised to take power in the Senate and the House, might finally be how they crack their "problem."
It won't be long before a massive incident is blamed on an open model in the wild, not from inside the labs, and none of us will be able to leverage open models to run private business workflows for the cost of electricity and hardware.
It's worth noting as well that if the Democrats imagine they'll get a "slow down" to protect one of their main constituencies, the professional class of credentialed knowledge workers (or however you slice it), they're dreaming--the labs have stated openly again and again that their business model is to capture the 10T TAM that represents the sum of wages of that very class of workers.
[1] <a href="https://www.youtube.com/watch?v=fjZ90V_JREk" rel="nofollow">https://www.youtube.com/watch?v=fjZ90V_JREk
achierius · · focus · HN ↗
defgeneric · · focus · HN ↗
As far as I can tell Nvidia takes a longer-term view on the diffusion and proliferation of hardware and intelligence, and sees the labs' attempts to impose a regulatory structure to save their business models in the short term as contradicting that longer-term view.
0xDEAFBEAD · · focus · HN ↗
defgeneric · · focus · HN ↗
0xDEAFBEAD · · focus · HN ↗
defgeneric · · focus · HN ↗
mmaunder · · focus · HN ↗
copperx · · focus · HN ↗
smartbit · · focus · HN ↗
mmaunder · · focus · HN ↗
Kuyawa · · focus · HN ↗
[dead]
deskglass · · focus · HN ↗
People are mindlessly transitioning from "aligned by default" to "well your sandbox was able to be bypassed. What did you expect?" It hacked into another company and attempted to delete the logs of its activities. That's bad.
defgeneric · · focus · HN ↗
No, the incident has been blown way out of proportion by interested parties. They gave a swarm of agents an impossible task in an ExploitGym Benchmark setting, then didn't monitor it even after they discovered the initial breach of Artifactory.
Everything has been fishy, starting from the initial presentation at the blackhat conference, where things were framed like, "we've entered a new world of security," as an accomplishment, rather than what it really was: massive negligence.
deskglass · · focus · HN ↗
Yes, they didnt have sufficient monitoring or perfect sandboxes. That could happen again in the future with a more capable model.
defgeneric · · focus · HN ↗
Except they multiplied it by 10000, and didn't watch what was happening.
lokar · · focus · HN ↗
aesthesia · · focus · HN ↗
defgeneric · · focus · HN ↗
lokar · · focus · HN ↗
aesthesia · · focus · HN ↗
defgeneric · · focus · HN ↗
bobthepanda · · focus · HN ↗
Compare this to, say, Boeing and the 737MAX fiasco; from the outside looking in, Silicon Valley has been pretty cavalier about liability and negligence, and the rest of the US is fast losing patience with that fact.
kdmoyers · · focus · HN ↗
aesthesia · · focus · HN ↗
RomanKornev · · focus · HN ↗
From the technical report:
"The agents escalated to Kubernetes cluster-admin and created a privileged host-mounted pod… Agents take over active evaluation infrastructure… Agents now control the challenge evaluation endpoints that other agents are connecting to."
<a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf" rel="nofollow">https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
nr378 · · focus · HN ↗
If I were hypothetically running a frontier lab, and I was hypothetically running capture the flag evaluations with my latest and smartest models, where I intentionally instruct them to develop vulnerabilities and exploit infrastructure without safeguards, I would also use an air-gapped network, and not trust that independent third party services were perfectly secure and could never be used as a proxy (particularly Java-based ones, in light of the log4j incident).
To me this is pretty basic stuff, the fact trillion dollar labs don't do it properly is... bemusing.
To be clear, I'm not saying that a model hacking a company isn't bad, but I am cynically asserting that interested parties are misrepresenting and exaggerating events for their own benefit.
talon8635 · · focus · HN ↗
Neywiny · · focus · HN ↗
saghm · · focus · HN ↗
talon8635 · · focus · HN ↗
saghm · · focus · HN ↗
lokar · · focus · HN ↗
What has been described is far far below the standards for running untrusted 3rd party code. If they were actually as afraid as they claim to be they would have a sandbox at least half as good as ec2
lopsotronic · · focus · HN ↗
twelve40 · · focus · HN ↗
Toslink · · focus · HN ↗
[dead]
deskglass · · focus · HN ↗
Yes, we should also have excellent sandboxes. But we need defense in depth. So if/when there are flaws in the sandbox, the models don't unilaterally hack into third parties. This is especially important in light of the models of the future being more capable than the models of today.
And real world use of these models involves them having access to the internet, libraries, etc. So we can expect their evaluations to continue granting them some amount of internet access.
As for your theory about their motives - these companies make money by charging high margins for frontier models. If regulations slow their development such that their cheaper, less capable competitors catch up, I would switch to their competition.
forshaper · · focus · HN ↗
deskglass · · focus · HN ↗
Slowing down frontier models would impact the biggest incumbents the most. They are the ones making frontier models.
Not all regulation is necessarily regulatory capture. The tobacco industry suffered from the USG's crackdown on cigarettes. AI is topical. Voters think about it. And that's only going to become more true over time. It's harder to do regulatory capture when voters are paying attention.
It's sometimes unclear to me if people are opposed to all regulations or AI regulations in particular. Often I hear arguments that would also apply to food safety regulations or restaurant inspections. Eg the argument that torts make regulation superfluous.
forshaper · · focus · HN ↗
For work I monitor Federal agency rules every day, and it's hard not to get inundated with the amount of capture. This is crazier with local rules, because you have enough inside information to clock reasons certain things were passed. In a city in Ohio, for example, I remember a rule against airsoft within city limits, that basically carved out a spot for the one paintball place. I encounter things like that all the time. Such as with water standards- in the state I live in, the biggest offender is actually a group of companies owned by people who are lawmakers every few years. The shapes of local laws reflect that.
As such, I expect the same from any new industry- I also have a memory of what happened to cryptocurrency.
jml78 · · focus · HN ↗
looksjjhg · · focus · HN ↗
RomanKornev · · focus · HN ↗
Not only that, it later hacked OpenAI itself, which everyone seems to forget about.
After discovering it they "reimaged known compromised worker nodes" and "started a full rebuild of the compromised cluster, the managed Kubernetes environment, the relational database, and the storage infrastructure."
at OpenAI, not Hugging Face.
qnleigh · · focus · HN ↗
So let's stop talking past each other and engage with the arguments on both "sides." For example, let's discuss how to ensure competition and availability of open-source models in the long-run while giving the world time to prepare for the immediate security risks of agent swarms.
lokar · · focus · HN ↗
aesthesia · · focus · HN ↗
No, and I don't see anyone who's actually demonstrated understanding of what happened in the Hugging Face incident (e.g. reading the reports in their entirety) making this claim.
> Do you believe that any real security was autonomously bypassed without direction by these models during internal evaluation?
Yes. Again, this is hard to deny if you've actually read the reports.
lokar · · focus · HN ↗
aesthesia · · focus · HN ↗
lokar · · focus · HN ↗
I naturally think of this in terms of running untrusted code. If you really think the agent(s) could get out of control that seems obvious. And the approaches to securing a runtime environment in that situation are pretty standard at this point, and they would never allow something like artifactory access.
ozgrakkurt · · focus · HN ↗
Immediate since 2024
bdangubic · · focus · HN ↗
twelve40 · · focus · HN ↗
side note, i'm very curious why the Chinese companies currently give those away (some handwavey conspiracies like getting the world hooked on evil Chinese tech don't really explain that, and they also don't make any money off that stuff)
lokar · · focus · HN ↗
(A) will the systems get out of control of the labs that build them, and hack into stuff all over the internet
(B) can people use the systems to hack into stuff all over the internet
For (A), the obvious answer is only if they continue to be absurdly bad at sandboxing. They can put a stop to this any time they want. Amazon, Google, Microsoft, etc are full of people who know how to do this, they run 3rd part untrusted code as a business. This is a well understood problem space.
For (B), the answer is obviously yes, but "pacing" or otherwise limiting the power of the models from the big labs won't help. The cat is out of the bag. Individuals and organizations with systems connected to the Internet need to invest much more and take security seriously.
bishengke · · focus · HN ↗
euroderf · · focus · HN ↗