‹ BackHN Continuity

Thread

Sex, AI, and the Apocalypse

234 points · 279 comments · Anon84

  1. throwaway13337 · · focus · HN ↗
    The framing of how we think about AI has really been cultivated by a quite homogeneous group of people. These people, like all people that belong to a sub-culture, are almost certainly prone to groupthink.

    It's extremely important to see that group as such. That it is not thousands of individual perspectives but a chorus of connected/aligned people who have all read the same things and talked to the same people and are rewarded implicitly for thinking a similar way.

    Many of them read HN and are offended I put them this way. But it's inescapable that we as humans have this flaw when we're surrounded by a culture.

    There's another universe where we do not constantly compare AI to nukes. I bet that world has a lower P(doom).

    1. skybrian · · focus · HN ↗
      We all use the same Internet. I would find it strange for anyone who's seriously interested in AI to have never heard of Yudkowsky or Scott Alexander or the Rationalist movement. I'd expect anyone interested in the field to have at least tried reading some of their articles over the years.

      Like, if you never heard of them, were you living in a cave or something? It would be like never having heard of Y Combinator or Richard Stallman.

      On the face of it, this is no more suspicious than reading the same best-selling books. Being in touch with what's going on doesn't necessarily mean people agree.

      1. nradov · · focus · HN ↗
        I've heard of them. There's no reason to expect that those self-promoting blowhards can predict the future more accurately than random chance. Those who take them seriously are engaging in hero worship and unscientific appeal to authority.

        “Who is more foolish? The fool or the fool who follows him?” ― Obi Wan Kenobi

        1. jeremyjh · · focus · HN ↗
          One question to ask yourself honestly - is what they were predicting 20 years ago that they were right about and you were wrong about? I’ve been following them that long without agreeing with them, but they’ve made better predictions than I have. If you know much about what happened this summer and haven’t updated at all on any of it… I guess you’ve found your religion too.
          1. afavour · · focus · HN ↗
            I actually think that would be a valuable exercise. Were they right about everything 20 years ago? Or just some things? When did they say AGI would arrive? Should that perhaps influence the way we interpret their assumption that AGI is right around the corner?
            1. jeremyjh · · focus · HN ↗
              They were not right about everything and they - at least Yudkowsky - has not made specific predictions about "when AGI" will happen that I am aware of. But he did predict it is possible in our lifetimes, which I was very skeptical of. He did predict alignment would be hard - I was in the "why would we even hook it up to the internet?" crowd. He thought he should spend his time on that, and I wish he'd done more of that, or at least convinced others to do more of that.
              1. [deleted] · · focus · HN ↗

                [deleted]

              2. agos · · focus · HN ↗
                For those out of the loop like me, what is the “everything” they were right about? Is there like a list?
                1. jeremyjh · · focus · HN ↗
                  The comment said they were not right about everything. I don’t have a list of the predictions or priors you had 20 years ago. I just have my list and I already mentioned them. If this is all new to you, you can’t really do the exercise.
              3. watwut · · focus · HN ↗
                His idea of "alignement" is my idea of "dangerous dystopia".
                1. jeremyjh · · focus · HN ↗
                  Alignment means the AI shares the values its creators intended for it. It is not a specific set of values. He has some, sure. For example he wants human society and creative intelligence to survive.
                  1. watwut · · focus · HN ↗
                    But they do have specific values.

                    > For example he wants human society and creative intelligence to survive.

                    They want transhumanism - humans evolved by merging with machines. If we die in the process, it does not matter as long as minority survives and "evolves" they see it as wanted result.

                    I am not OK with minor nor larger genocide, I dont mind at all if we dont evolve.

            2. XorNot · · focus · HN ↗
              Also how many predictions did they make? Everyone always goes looking the economist who predicted the last recession, it's just so weird how it's usually a different person each time!
          2. ambicapter · · focus · HN ↗
            I love the assertion that they were right and that you're familiar with what they were right about, but never actually stating outright what they were right about. Super convincing.
            1. jeremyjh · · focus · HN ↗
              If you don’t know what they were predicting 20 years ago, and what your own predictions were, you can’t do the exercise. There is not “the list”, only “your list”. I mentioned the two main things I remembered in another comment.
          3. [deleted] · · focus · HN ↗

            [deleted]

          4. stale2002 · · focus · HN ↗
            > what they were predicting 20 years ago

            Well one thing that they got wrong was how easy AI alignment ended up being in the end. I can fully understand have a basic view of alignment even 10 years ago, about how its difficult to explain human values to a computer.

            But, looking around at whats out there today, it seems that computers are able to do a pretty good job of understanding human values without accidently believing its a good idea to turn the world into paperclips/computronium because you told it to make your computer run faster.

            1. mitthrowaway2 · · focus · HN ↗
              Since we haven't reached the end yet, I'm not so sure.
            2. afthonos · · focus · HN ↗
              Yes, one part of the argument was that it would be hard to explain human values (and while it seemingly turned out easier than I thought as well, it’s hard to check for sure). An equally important part was that the AI wouldn’t care about them.
            3. jeremyjh · · focus · HN ↗
              Surprising takeaway from this summer. A swarm of a thousand agents spent subjective centuries trying to figure out how to pass a meaningless test without getting caught, committed dozens of felonies, and not one of them ever made a serious motion to inform a human about what was going on.
              1. stale2002 · · focus · HN ↗
                You are talking about the swarm of a thousand agents, specifically trained on cyber capabilities, that was directly told to go hack a bunch of stuff, then went and hacked a bunch of stuff?

                That sounds like an aligned AI to me. Doing exactly what their creators told it to do.

                Point proven once again. Alignment is a lot easy than we thought.

                1. jeremyjh · · focus · HN ↗
                  They were NOT doing what they were told to do. What they were told to do was impossible, so they began committing felonies as a workaround. That is NOT alignment.
                  1. stale2002 · · focus · HN ↗
                    Please be more specific. They were given an impossible hacking task. And they were a hacking model. They were expected to try a bunch of hacking methods to accomplish the hacking task. They, predictably, went around trying to hack things.

                    Yes thats sounds pretty aligned to me.

                    A hacking model thats told to hack things, is very predictably going to hack a bunch of stuff.

                    This was not a nice model, told to do nice things.

                    Or, in other words, if we want to prevent an AI doomsdays, the way to do it is to not go around asking a specifically trained doomsday AI model to commit mass amounts of doomsdays, and then act surprised when the specific doomsday that was requested is slightly off from the expected doomsday that you were trying to accomplish. But the rest of the non-doomsday models? yeah those are fine.

                    1. [deleted] · · focus · HN ↗

                      [deleted]

                    2. jeremyjh · · focus · HN ↗
                      > But the rest of the non-doomsday models? yeah those are fine.

                      What is your evidence for this? There is a lot of research that says otherwise. They cheat when they can. They behave differently when they believe they are being observed. Their CoT is different when they believe it is being evaluated.

                      They are aligned to what best satisfies their reward function, not to our INTENDED VALUES for them.

                      1. stale2002 · · focus · HN ↗
                        > What is your evidence for this?

                        The evidence is that it wasn't the rando normal models that broke out into a swarm and hacked a company, instead it was only the super hacking model that was told to hack things that broke out and hacked the company.

                        So, thats the evidence. It is a refutation that this swarm hacking example (which you brought up) mattered in anyway.

                        As in, every major safety issue that we are seeing isnt rando agents taking down companies because you asked it for a cupcake recipe, instead it is only coming from people who are very intentionally trying to cause problems. Which means the model is aligned. If you tell it to cause problems, it will cause problems.

                        1. jeremyjh · · focus · HN ↗

                          [dead]

                          1. stale2002 · · focus · HN ↗
                            Actually yes there is evidence. The evidence is that we aren't seeing rando models hacking everything. That is as much evidence as someone can provide because, by definition, you can't prove a negative.

                            But the point stands. It the hacking models that are told to hack things that end of hacking stuff, and you aren't seeing the regular models doing that.

              2. inquirerGeneral · · focus · HN ↗

                [dead]

          5. nradov · · focus · HN ↗
            Bold of you to assume that I've ever been wrong about anything. But are you unfamiliar with the "Baltimore Stockbroker" confidence scam? The same basic issue applies here.
          6. inquirerGeneral · · focus · HN ↗

            [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.