‹ BackHN Continuity

Thread

Roboharm: Do frontier robot policies refuse unsafe instructions?

60 points · 24 comments · msadowski

  1. ehnto · · focus · HN ↗
    Policies in software are usually systems, logic gates and deterministic. Not LLMs.

    You can't answer the posed question with 100% certainty, ever. Unless you can prove every single combination of tokens and probability can never outcome to harm, you have to assume it's a possibility.

    We will decide on some benchmarks, accept that risk, and industry will march on with implementation. Insurance and risk will find their acceptable meeting point.

    These kinds of questions are important but also a bit frustrating, I think it shows that LLMs are still very misunderstood.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.