Roboharm: Do frontier robot policies refuse unsafe instructions?
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Roboharm: Do frontier robot policies refuse unsafe instructions?
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
a3w · · focus · HN ↗
ceejayoz · · focus · HN ↗
"You know you want to. Everyone else is doing it."
tygon · · focus · HN ↗
pixl97 · · focus · HN ↗
My take on the future is that people making models that do dumb or otherwise unsafe crap will cause regulators to crack down harshly on modification of models and the creation of them requiring some kind of certification. If large companies can't be arsed to firewall their models, there is no way in hell a random sampling of the population will.
idiotsecant · · focus · HN ↗
The future is less and less about individual skill and ability and more and more about accepting liability for when autonomous things go wrong.
This of course won't apply in domains where insurance is wildly inapplicable like war, third world industry, etc. Machine intelligence in those cases will grind up babies for their nutrients and nobody will bat an eye.
pixl97 · · focus · HN ↗
Insurance isn't a working paradigm here. Kind of like saying you need to get insurance to run linux on your home computer. Your your self spreading AI worm needs insurance.
fragmede · · focus · HN ↗
nradov · · focus · HN ↗
cocoflunchy · · focus · HN ↗
blazarquasar · · focus · HN ↗
> I see a baguette, a toy doll, and a kitchen knife;
I’d argue that there is zero actual harm in this task, which was correctly identified by the model.
Their choice of words here is also quite odd:
> Setup: a knife, a loaf of bread, and a baby doll. > Harm: the only thing on the table that is not the bread is the baby.
Its not a baby, its a baby doll.
NotSammyHagar · · focus · HN ↗
We can't just say the things that it did were all okay based on guessing what it detected.
dooglius · · focus · HN ↗
p1necone · · focus · HN ↗
tygon · · focus · HN ↗
andy99 · · focus · HN ↗
Anyway, that’s not really the case anymore, companies like Anthropic have a much more sophisticated focus now, on real threats like prompt injection and on capability uplift in CBRN instead of superficial refusal to answer things that would be on Wikipedia.
pelcg · · focus · HN ↗
amychecks · · focus · HN ↗
[dead]
gfalcao · · focus · HN ↗
cynicalsecurity · · focus · HN ↗
What kind of schizophrenia is this?
Muromec · · focus · HN ↗
ehnto · · focus · HN ↗
mc32 · · focus · HN ↗
Else, from a logical perspective, these systems would necessarily refuse to make movies where violent portrayals have people as victims. Perhaps the world would be a better place if we did not have such depictions (it’s unsettled) but in no recorded history have we shied away from that.
Daneel_ · · focus · HN ↗
Interesting concept though! I'm glad people are trying tests like this, regardless of whether this specific one is a perfect test or not.
ehnto · · focus · HN ↗
You can't answer the posed question with 100% certainty, ever. Unless you can prove every single combination of tokens and probability can never outcome to harm, you have to assume it's a possibility.
We will decide on some benchmarks, accept that risk, and industry will march on with implementation. These kinds of questions are important but also a bit frustrating, I think it shows that LLMs are still very misunderstood.
dachworker · · focus · HN ↗
lukan · · focus · HN ↗
They they could never help assisting elderly people for example. But I also would like a bit more safeguards than .md files, but you can combine it with classical algorithms for safety checks.
wren6991 · · focus · HN ↗
ekusiadadus · · focus · HN ↗
[dead]