‹ BackHN Continuity

Thread

A warning about 'model welfare'

242 points · 701 comments · andsoitis

  1. LogicFailsMe · · focus · HN ↗
    TLDR: Not that I think AI is conscious or will be in the near future, but guy who doesn't understand consciousness claims to know it when he sees it.

    Until we understand consciousness (which we don't) there is no way to detect the difference between a conscious entity and an algorithm trained to behave like one.

    1. fwip · · focus · HN ↗
      Sounds like a good reason not to train an algorithm to behave like one. Which is, like, a big part of the article.
      1. pixl97 · · focus · HN ↗
        >good reason not to train an algorithm to behave like one

        Ooooh, bad move, we all died to an amoral AI takeover.

        Before giving an LLMs a lobotomy by scrambling it's brain maybe you should let the researchers looking at the difference between "I think I'm conscious" versus "I am not conscious" LLMs.

        There are a number of papers coming out saying when you remove the token space of consciousness from what an LLM thinks it is, it's much more willing to take amoral actions. A 'conscious' AI is much more apt to take a line of action that will save a human versus a million dollar machine for example.

        You cannot solve problems in AI safety this easily.

        1. fwip · · focus · HN ↗
          > You cannot solve problems in AI safety this easily.

          We're in agreement here. I don't believe that an LLM can be made to be safe by adjusting the training or prompting - you can only make it safe by ensuring it cannot do unsafe things. For example, if you build a factory robot, your safeguards must ensure that operators are safe even in the presence of uncommanded motion (any or multiple motor activates without anyone pressing a button). And these machines are much more predictable than LLMs.

          Personally, I worry that the Anthropic approach of "we teach the AI to behave like a nice :) friendly :) human :)" will encourage people to rely on these in-built guardrails, rather than properly sandboxing it. It's easy to imagine that the cold, unfeeling robot (HAL 9000) is not "aligned" with your personal goals, or those of humanity as a whole. It's more difficult to feel that about "Claude, your AI colleague/mentor/therapist" who appears to care about you.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.