‹ BackHN Continuity

Thread

US Military had close call after using AI for hallucinated intelligence report

519 points · 396 comments · realsarm

  1. drtgh · · focus · HN ↗
    > relatively poorly understood technology

    Poorly understood? how convenient...

    LLMs are vectorial databases with losses that index statistically filled data, which uses a text interface to query such statistically filled data. The output is a string concatenation (statistically concatenated bit by bit).

    When the LLMs are queried (prompted), you can get random mixed data as output, ERRORS, due to undesired indexes getting closer at one point while the string was being concatenated for the output, what affects the rest of the indexed content that will be concatenated.

    It is intrinsic to this tech. The larger the context, the greater the probability of get mixed data. And if the provider lowers the precision of those indexes -in order to decrease hardware resources and energy consumption- such probability increases to the point where those errors are granted.

    Even knowing that the queries can return wrong/mixed data in the responses, errors, the companies developing this, decided to introduce a new product, that connects such LLMs outputs to the command console, latter connected to internet, raw 'eval' running commands from such outputs witch obviously can contain whatever mixed random. Then we started to hear "oh, it deleted my directory", etc, and it seems the next one will be "a missile killed my wife", because it is a text concatenation engine with errors.

    To name it "hallucination" is an euphemism... those are errors, and they are granted to happen at one moment. If they do not know this, then they ate too much marketing without doing their job, or it was a convenient contract for the pocket$ of someone.

    1. margalabargala · · focus · HN ↗
      I agree with most of your comment, but...

      > To name it "hallucination" is an euphemism... those are errors

      I find this and other "don't anthropomorphize the computer" statements incredibly unconvincing.

      People develop terms for things and language has always contained overloaded or "literally inaccurate" terms.

      An LLM can have "hallucinations" in the same way a modern computer program can have "bugs".

      1. 0x20cowboy · · focus · HN ↗
        It’s not an error or a hallucinations it works correctly every time, and statistically picks the next token for the sequence.

        Retuning inf or crashing would be an error.

        If you want to ascribe some kind of meaning to the tokens, then maybe the training data was insufficient to predict the token in the sequence you wanted, but it doesn’t predict the next “fact”, and it doesn’t “think” it predicts the next token.

        1. margalabargala · · focus · HN ↗
          LLMs are useful because (and inasmuch as) their output generally reflects coherent reality.

          And their output does, usually, reflect coherent reality.

          The problem class of "properly operating program emits output incompatible with coherent reality" is something that is reasonable to put under its own term, considering it's a new class of problem.

          In other words, I think you misunderstand the language others are using. "Hallucination" doesn't refer to an "error" in the sense that crashing is an error, it refers to a situation in the problem class above, which is compatible with it working correctly every time.

          > it doesn’t “think” it predicts the next token.

          I never said it did. And I agree that LLMs don't "think". That said I am fully willing to go to bat arguing "thinking tokens" is a perfectly fine piece of jargon. Metaphors are completely acceptable parts of language, and contextual meaning is something grasped by everyone including the pedants who pretend not to.

          1. 0x20cowboy · · focus · HN ↗
            > I think you misunderstand the language others are using. "Hallucination" doesn't refer to an "error" in the sense that crashing is an error, it refers to a situation in the problem class above, which is compatible with it working correctly every time.

            I do not misunderstand, I think maybe you do. You think there is a proper next word selection based on logic or meaning and there for the model selected the wrong one - it hallucinated.

            I am saying the model has no concept if anything other than the probability of select a token which is not based in any logic so it is working properly- it only works on numbers.

            It is random chance that it is ever correct, not that it is correct often and messed up this one time.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.