‹ BackHN Continuity

Thread

US Military had close call after using AI for hallucinated intelligence report

519 points · 396 comments · realsarm

  1. drtgh · · focus · HN ↗
    > relatively poorly understood technology

    Poorly understood? how convenient...

    LLMs are vectorial databases with losses that index statistically filled data, which uses a text interface to query such statistically filled data. The output is a string concatenation (statistically concatenated bit by bit).

    When the LLMs are queried (prompted), you can get random mixed data as output, ERRORS, due to undesired indexes getting closer at one point while the string was being concatenated for the output, what affects the rest of the indexed content that will be concatenated.

    It is intrinsic to this tech. The larger the context, the greater the probability of get mixed data. And if the provider lowers the precision of those indexes -in order to decrease hardware resources and energy consumption- such probability increases to the point where those errors are granted.

    Even knowing that the queries can return wrong/mixed data in the responses, errors, the companies developing this, decided to introduce a new product, that connects such LLMs outputs to the command console, latter connected to internet, raw 'eval' running commands from such outputs witch obviously can contain whatever mixed random. Then we started to hear "oh, it deleted my directory", etc, and it seems the next one will be "a missile killed my wife", because it is a text concatenation engine with errors.

    To name it "hallucination" is an euphemism... those are errors, and they are granted to happen at one moment. If they do not know this, then they ate too much marketing without doing their job, or it was a convenient contract for the pocket$ of someone.

    1. theptip · · focus · HN ↗
      > LLMs are vectorial databases

      You use a bunch of technical-sounding words here to make it sound like you understand. But to be clear, nobody understands why the evolved weights of a NN make the decisions that they do.

      Almost nothing is understood about the actual representations used for nontrivial concepts, decision algorithms, etc.

      If you look at the field of mechanistic interpretability, compared to “GOFAI” like learned decision trees, an LLM is completely opaque.

      1. Betelbuddy · · focus · HN ↗
        >> an LLM is completely opaque

        And despite that, although they are not like that in practice as there are too many uncontrolled variables, with temperature at zero, for the same input they produce always the same reply.

        1. chrisjj · · focus · HN ↗
          > with temperature at zero, for the same input they produce always the same reply.

          Nonsense.

          <a href="https:&#x2F;&#x2F;thinkingmachines.ai&#x2F;blog&#x2F;defeating-nondeterminism-in-llm-inference&#x2F;" rel="nofollow">https:&#x2F;&#x2F;thinkingmachines.ai&#x2F;blog&#x2F;defeating-nondeterminism-in...

          1. coldtea · · focus · HN ↗
            BS. Run them sequentially on a single core, and without fancy speedups enabled, and they do. The algorithm is determinstic. Any non-determinism present with 0 temperature it&#x27;s not some mysterious LLM-inherent property, but something that can be seen in any large program taking advantage of multi-core, floating point, and other CPU-based parallelism optimization.
            1. chrisjj · · focus · HN ↗
              &gt; Any non-determinism present with 0 temperature it&#x27;s not some mysterious LLM-inherent property,

              Of course it&#x27;s not mysterious. It is well understood by all who are aware of the fundamental unreliability of all major LLMs in general use today.

              1. coldtea · · focus · HN ↗
                The phrasing still makes it sound like it&#x27;s due to the LLM. It&#x27;s not. It&#x27;s do to compiler and CPU, and OS optimizations, and a non-LLM program could suffer the same just as well.
                1. chrisjj · · focus · HN ↗
                  This unreliability is entirely due to the implentation of the LLM. Yes, any other program implemented equally carelessly could suffer the same. But you&#x27;d be hard pressed to find any as unreliable as a typical LLM.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.