‹ BackHN Continuity

Thread

A warning about 'model welfare'

242 points · 701 comments · andsoitis

  1. catigula · · focus · HN ↗
    I largely agree that this claim is likely correct, but as far as I understand the science, this specific claim;

    >They do not have innate preferences or underlying motivations

    Is incorrect unless you’re being extremely pedantic in an intellectually unhelpful way.

    1. pixl97 · · focus · HN ↗
      The crazy thing is agents with long running memory/context do have preferences that become individualized. Almost every time I see a detractor around these things they typically have large gaps of what some people 'growing' these systems to do.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.