‹ BackHN Continuity

Thread

Our framework for reporting model misalignment

107 points · 96 comments · qprofyeh

  1. thewhitetulip · · focus · HN ↗
    If model labs can't control astra level model, how can they control AGI?!

    Seems like there are no guardrails on LLMs

    1. worldsavior · · focus · HN ↗
      No one can control any AI model. It will never be controlled. These models are based on a huge amount of data, it's just gonna be impossible to control the output that is based on that data only with a system prompt or some other injection mechanism.
      1. dns_snek · · focus · HN ↗
        The model is just a powerless token generator without a harness. If you give the model a harness which you choose to exercise no control over, can you say that it can't be controlled?
        1. pizza234 · · focus · HN ↗
          Inform yourself by reading the METR analysis of the HuggingFace incident.

          Agents simply broke out of their environment. And this can't be discarded anymore by assuming that it's just a poorly configurend jail, because agents are becoming better and better at escaping.

          In short: on a large enough scale and timeline, the possibility of constrain AIs approaches zero.

          Bonus: what many people don't know is that agents also hacked in the internal OpenAI network. Crazy times.

          1. timr · · focus · HN ↗
            While informing yourself, don't skip the part where you find out that "the environment" was the security equivalent of a wet paper bag.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.