‹ BackHN Continuity

Thread

Our framework for reporting model misalignment

107 points · 96 comments · qprofyeh

  1. thewhitetulip · · focus · HN ↗
    If model labs can't control astra level model, how can they control AGI?!

    Seems like there are no guardrails on LLMs

    1. worldsavior · · focus · HN ↗
      No one can control any AI model. It will never be controlled. These models are based on a huge amount of data, it's just gonna be impossible to control the output that is based on that data only with a system prompt or some other injection mechanism.
      1. dns_snek · · focus · HN ↗
        The model is just a powerless token generator without a harness. If you give the model a harness which you choose to exercise no control over, can you say that it can't be controlled?
        1. pizza234 · · focus · HN ↗
          Inform yourself by reading the METR analysis of the HuggingFace incident.

          Agents simply broke out of their environment. And this can't be discarded anymore by assuming that it's just a poorly configurend jail, because agents are becoming better and better at escaping.

          In short: on a large enough scale and timeline, the possibility of constrain AIs approaches zero.

          Bonus: what many people don't know is that agents also hacked in the internal OpenAI network. Crazy times.

          1. dns_snek · · focus · HN ↗
            1. You misunderstood my comment. Models can't escape, they can't do anything, they only generate tokens. Models become agents when you add a harness which is simultaneously a leash around the model.

            The model merely requests that your harness do something. If your harness just executes every request without oversight then you can hardly complain when it does something unintended.

            This is foundational, we're not even talking about the OS/network-level sandboxing that should be applied on top of this.

            2. Like another comment already pointed out, that sandbox OpenAI used was the equivalent of a wet paper bag. Artifactory is not meant to be a security boundary for malicious payloads.

            1. citrin_ru · · focus · HN ↗
              It's so tempting (because it's valuable) to give a model access to the internet (via harness) that the only way to stop people from doing this is some enforceable legislation or stricter liability when people will not be able to avoid responsibility by saying it's not me, it's AI on it's own.
              1. ozgung · · focus · HN ↗
                OpenAI case was actually an exception. Agents had no internet access because they were evaluated for a benchmark. In real life agents have access to virtually everything, most people use them like that.

                If any of you actually know how to make agents secure (without limiting everything) you can be a billionaire.

                1. dns_snek · · focus · HN ↗
                  Sure, we can make them very secure at the cost of some convenience by enforcing narrowly scoped permissions/capabilities for everything with human approval. But that takes some effort so people actually want YOLO mode without any trade-offs or risks which is impossible.

                  If you decide to do it anyway then you bear the consequences of those decisions. First comparison that comes to mind is driving drunk and hitting someone.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.