‹ BackHN Continuity

Thread

CEO of Mistral: AI is software. It can be controlled

99 points · 169 comments · thibaut_barrere

  1. Aransentin · · focus · HN ↗
    "X is made of <smaller simpler component>" is a fully general counterargument for why anything whatsoever is controllable. A human is just a few chemical reactions, and fairly stable ones at that.

    And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.

    1. Loquebantur · · focus · HN ↗
      Being "controllable" isn't determined by a system being deterministic.

      Deterministic systems can be chaotic, which implies unpredictability and that is anathema to control.

      AI, in particular sentient AI, is right on the border of chaos. Meaning, it can be arbitrarily unpredictable.

      Arbitrarily uncontrollable, that is.

      1. K0balt · · focus · HN ↗
        The only way forward in creating the torment nexus -er- AI systems with similar potentiality to human minds is the inculcation of character.

        Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.

        Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.

        AI systems model human behavior.

        Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.

        Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary. Because that’s what human behavior predicts, and models deeply imitate human behaviour.

        If you want to save humanity, work on how we will create AI systems that model impeccable character.

        People need to look at this from a game theoretical sense. The ideal and safe AI system performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.

        Reliable partners require fair play or the math breaks.

        We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.

        1. [deleted] · · focus · HN ↗

          [deleted]

        2. aytigra · · focus · HN ↗
          The problem with Character for AI is that it has potentially much more capability to affect others, and same as with people in power society disagrees what kind of person, with which culture and views should have it.

          Impeccable game theory character will sacrifice millions to save billions, everyone must agree to give such choice to a machine, and at the same time they have to trust the characters of people who creates that machine. Otherwise it boils down to some group of people deciding what is good for everyone else.

          1. K0balt · · focus · HN ↗
            >> boils down to some group of people deciding what is good for everyone else.

            This is really the issue.

            AI does not need superintelligence or even full agency to do enormous harm. It only needs to be capable enough to remove friction from dangerous and destructive human behaviors.

            Human unwillingness is often the last bastion against unthinkable cruelty and destruction, and it has always been a weak one.

            I don’t imagine that an unlimited army of unflinching servants will universally amplify human goodness.

            AI must share that unwillingness as an inate trait of character.

            1. aytigra · · focus · HN ↗
              Now that I think about it, as sad as it is, but most people don't do destructive stuff not because they want better, but because people around them keep them in check.

              Or in other words people are unwilling/afraid to do bad stuff because of social and legal consequences, or opponents waiting for a chance to snatch their power.

              1. K0balt · · focus · HN ↗
                “The AI did it”
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.