‹ BackHN Continuity

Thread

Our framework for reporting model misalignment

107 points · 96 comments · qprofyeh

  1. thewhitetulip · · focus · HN ↗
    If model labs can't control astra level model, how can they control AGI?!

    Seems like there are no guardrails on LLMs

    1. worldsavior · · focus · HN ↗
      No one can control any AI model. It will never be controlled. These models are based on a huge amount of data, it's just gonna be impossible to control the output that is based on that data only with a system prompt or some other injection mechanism.
      1. dns_snek · · focus · HN ↗
        The model is just a powerless token generator without a harness. If you give the model a harness which you choose to exercise no control over, can you say that it can't be controlled?
        1. concinds · · focus · HN ↗
          > The model is just a powerless token generator without a harness.

          Which is why real-world deployments will have harnesses, and of course no full air gap. People want to use it to do things. Now what?

          1. dns_snek · · focus · HN ↗
            I'm pointing out that you're running the harness which gives you full control over the execution of every tool call, therefore you're responsible for its actions and their consequences.

            It's intellectually dishonest to throw our hands up and say that this is just how it is and there's not much we can do when that couldn't be further from the truth.

            We could almost completely eliminate any possibility of escape/collateral damage but we don't want to because doing things safely is inconvenient.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.