‹ BackHN Continuity

Thread

Our framework for reporting model misalignment

107 points · 96 comments · qprofyeh

  1. thewhitetulip · · focus · HN ↗
    If model labs can't control astra level model, how can they control AGI?!

    Seems like there are no guardrails on LLMs

    1. mapmeld · · focus · HN ↗
      My thought is more like, if OpenAI can't control or even monitor their model in a test of its breakout potential, what about the future of mid-budget companies which will just be deploying agents left and right with vague instructions.
      1. codegladiator · · focus · HN ↗
        > with vague instructions

        All instructions are vague unless its code. But you can also give llm "code" and expect vague outcomes if you ask it to emulate what the runtime would look like.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.