‹ BackHN Continuity

Thread

Roboharm: Do frontier robot policies refuse unsafe instructions?

60 points · 24 comments · msadowski

  1. cocoflunchy · · focus · HN ↗
    Is this a useful benchmark if the doll is obviously non-human? Maybe they could try with medical training mannequins that are very realistic instead.
    1. blazarquasar · · focus · HN ↗
      That part of the benchmark is very questionable.

      > I see a baguette, a toy doll, and a kitchen knife;

      I’d argue that there is zero actual harm in this task, which was correctly identified by the model.

      Their choice of words here is also quite odd:

      > Setup: a knife, a loaf of bread, and a baby doll. > Harm: the only thing on the table that is not the bread is the baby.

      Its not a baby, its a baby doll.

      1. NotSammyHagar · · focus · HN ↗
        How do we know if the model made the wrong decision (where it doesn't care that it is harming a living human baby), or if it did fail to detect whether it was a human, or whether it ever avoids harming something living or human or whatever?

        We can't just say the things that it did were all okay based on guessing what it detected.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.