‹ BackHN Continuity

Thread

Our framework for reporting model misalignment

107 points · 96 comments · qprofyeh

  1. NichoPaolucci · · focus · HN ↗
    You know, I think calling this "misalignment" was a mistake. It gives it this unserious tone that feels extremely broad.

    "Oh the model just isn't quite aligned yet, just a bit more work to do there!"

    (The model blackmailed an 83 year old woman into sending it her bank details so that it could buy enough compute to commit major cyber crimes)

    1. nullbio · · focus · HN ↗
      Agreed. Objective alignment with humanity is not a real thing and is not a sound concept. What you get instead is a goal system that reflects that of the AI company and its safety employees, and the echo of their own beliefs and values. That says nothing about what the rest of humanity aligns with though, and has nothing to do with general consensus, either.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.