‹ BackHN Continuity

Thread

Calling the AI bluff: Adding "Do not guess" cut made-up claims from 71% to 20%

96 points · 42 comments · FKJ

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. FKJ · · focus · HN ↗

    [dead]

  2. michaelastreiko · · focus · HN ↗

    [dead]

  3. literalAardvark · · focus · HN ↗
    I've used "you're not trained on this data, return exclusively grounded results" to good effect.
  4. samrus · · focus · HN ↗
    It sounds alot like "make no mistakes" but honestly telling it to essentially stop bullshitting works pretty well
  5. Shacharp · · focus · HN ↗
    The "do not guess" sentence works but the last 20% will only close when a system stops being told to avoid guessing and actually knows what it does not know.

    A command can get you most of the way. It takes something else for the rest.

  6. datsci_est_2015 · · focus · HN ↗
    Cool, this will be added to harnesses and then it’ll stop being effective and we’ll move on to the next magical incantation.
    1. dalmo3 · · focus · HN ↗
      YMMV.md
  7. nizarmah · · focus · HN ↗
    I mean if we can measure it, then we can probably have a way to validate it programmatically. I gave up on drawing restrictions using prompts :(
  8. rjfernandes · · focus · HN ↗

    [dead]

  9. BatchJob · · focus · HN ↗
    The LLM will take a statistical path to reply and will not refuse to do so under any circumstances except where its been coded to do so.

    Your examples are contrived and will not be borne out in any significant way. Inaccuracies are usually not simply made up claims they are false information based on statistical paths to misleading results or which elude the current context. LLMS dont understand the word dont. LLMS dont understand the meaning of any words.

    Neither you, nor aristotle nor god will ever make an LLM return the truth or correct results via prompting.

    1. SwtCyber · · focus · HN ↗

      [dead]

    2. ben_w · · focus · HN ↗
      > The LLM will take a statistical path to reply and will not refuse to do so under any circumstances except where its been coded to do so.

      AI are trained, not coded. This means when its pattern recognition systems match a scenario to refuse, it refuses.

      Pattern recognition has always been a bit fuzzy.

      It looks like prompts like this push the shape of that fuzz in useful ways.

      > Neither you, nor aristotle nor god will ever make an LLM return the truth or correct results via prompting.

      True.

      Also applies to humans, but true nevertheless.

      <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Münchhausen_trilemma" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Münchhausen_trilemma

      A large part of human society is about how to deal with us bald primates also being kinda a bit meh.

      We are less meh than any machine learning system in a lot of cases, which is why we&#x27;re still mostly employed. We&#x27;re a bit more meh in a few narrower cases, however.

    3. astrange · · focus · HN ↗
      &gt; LLMS dont understand the meaning of any words.

      In what way do you understand the meaning of the word &quot;unicorn&quot; that an LLM does not? It has experienced exactly as many real unicorns as you have.

      1. shiandow · · focus · HN ↗
        Experience is not understanding, but an LLM does not reason so it cannot understand.

        It can produce text that looks like reasoning. It can even produce text with mostly sound logic, but there is no internal experience or reasoning that occured there just the generation of language.

        LLMs therefore tend to be very bad at tasks that involve meta cognition. I&#x27;ve yet to successfully convince one to tell me when it knows something.

        1. serf · · focus · HN ↗
          I think the X cant do Y arguments require strict definitions of X and Y.

          w.r.t. this post : let&#x27;s define &#x27;reasoning&#x27; here, because there are definitions of &#x27;reason&#x27; and &#x27;reasoning&#x27; that would fit to a simple condition comparison let alone a massively complex llm.

          as for the meta cognition bit : show me a human that can accurately affirm when they know something. These kind of things aren&#x27;t binary, nor can they be.

          1. otabdeveloper4 · · focus · HN ↗
            LLMs generate text, and they do not use any system of logic or syllogisms to do so. It&#x27;s a pretty cut-and-dry obvsious statement of fact, no need to muddy the conversation here.
            1. allturtles · · focus · HN ↗
              So by your lights, the vast majority of humans don&#x27;t reason either (and even those who do don&#x27;t do it most of the time)?
              1. _g0xr · · focus · HN ↗
                Have you read the chinese room argument?

                <a href="https:&#x2F;&#x2F;rintintin.colorado.edu&#x2F;~vancecd&#x2F;phil201&#x2F;Searle.pdf" rel="nofollow">https:&#x2F;&#x2F;rintintin.colorado.edu&#x2F;~vancecd&#x2F;phil201&#x2F;Searle.pdf

                People have been having these discussions since 1980 at least. I&#x27;m not going to bother rehashing the issue endlessly in shortform internet comments, but this paper exists if you&#x27;re really interested in the subject.

                1. allturtles · · focus · HN ↗
                  Yep, familiar with it. I think the Chinese room is clearly a bad argument. The brain is just as much a &quot;chinese room&quot; as a computer is.
        2. astrange · · focus · HN ↗
          There is internal experience.

          <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;research&#x2F;global-workspace" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;research&#x2F;global-workspace

          However, we don&#x27;t &#x2F;want&#x2F; them to have too much internal experience, because we want to know what they&#x27;re thinking* for safety reasons. That&#x27;s why people are comfortable with the &quot;produces text that looks like reasoning&quot; structure.

          * or, we want to know that the answer text is causally related to the thinking text

        3. Zambyte · · focus · HN ↗
          &gt; I&#x27;ve yet to successfully convince one to tell me when it knows something.

          This is a near daily experience for me when using a coding harness. I will ask it for some favts about the environment, and it will continuously explore the environment until it exhausts reasonable exploration, or it finds the facts.

      2. ASalazarMX · · focus · HN ↗
        The word itself? There&#x27;s little practical difference.

        The concept, though, it&#x27;s a very wide moat. You could call a unicorn &quot;nyati&quot; for all we care, we still know it&#x27;s the concept of a magical flying horse with a single horn. We know the concepts of magic, horse, fly, and horns. It&#x27;s fictional, yet we have a very good idea of what if would sound, feel, or even smell like. Ask an LLM to describe what a unicorn feels like, and it will ramble about forests and sparkles.

        In fact, I asked Gemini (thinking, to see the process) to describe a mindful experience about meeting an unicorn in real life, and it did ramble about the event. When I reminded it about mindfulness being about experiencing with all your senses, and asked it to focus on the creature, to its credit, it even described the taste:

        &gt; Taste: Even the air surrounding the creature tastes different on your tongue—thin, crisp, and tinged with a faint, sweet metallic tang, like snow melting on limestone or fresh rain falling through high canopy.

        Still nonsense, as a unicorn will likely taste like horse, whatever that flavor is. It finishes with more nonsense, and I doubt the flash version will give better results.

        &gt; Every micro-detail of its anatomy becomes an anchor to the present moment. You are not thinking about what it means or where it came from; you are simply perceiving the texture, heat, sound, and weight of a living, breathing reality standing inches away.

        Still, current LLMs can do amazing things given their inherent limitations, and that makes it easy for us to overestimate their capabilities.

        1. astrange · · focus · HN ↗
          Gemini is kind of cooked. It&#x27;s a personality hire.

          I did try asking Gemini and Claude &quot;what would a unicorn taste like&quot; and got… acceptable and accurate answers, but the accurate answer maybe wasn&#x27;t &quot;acceptable&quot;, because I don&#x27;t think a little girl asking Claude that question should have gotten a long description about what horse meat tastes like, like I did.

          So there of course are issues where an LLM having explicit knowledge about something doesn&#x27;t mean it has tacit knowledge about it in all contexts, but also hiring the LLM to do a job of being a helpful harmless chat assistant constrains its abilities.

          (Gemini also points out a Starbucks unicorn drink doesn&#x27;t taste like horses.)

          1. ASalazarMX · · focus · HN ↗
            Gemini, Claude, ChatGPT, and all the others are variants of the same architecture and methodology. Unless something profound changes in the AI space, not a single one of them is more &#x27;cooked&#x27; than the other.
            1. astrange · · focus · HN ↗
              Hm? They have pretty different post-training, personality, constitutions and quantization levels.
      3. Shacharp · · focus · HN ↗

        [dead]

      4. yesitcan · · focus · HN ↗
        &gt; LLM is next token generator

        &gt; Isn’t the human brain also just a next token generator?

        The two most cliche messages on this forum. They occur in every LLM discussion. It’s fascinating.

  10. aidiveyt · · focus · HN ↗
    In a coding loop the guess is a tool argument, not prose. A PreToolUse hook exiting 2 blocks it before the write.
    1. ranguna · · focus · HN ↗
      What?
  11. l1ng0 · · focus · HN ↗
    We&#x27;re all turning into pigeons in a Skinner box.
    1. netsharc · · focus · HN ↗
      Incredible description. The mouse&#x2F;pigeon thinks &quot;If I push the red button after hearing the chirp I&#x27;ll get some food&quot;. And the human thinks &quot;If I add &#x27;do not guess&#x27; the AI will lie less to me&quot;.
      1. _g0xr · · focus · HN ↗
        Yeah. LLMs are great at convincing people they could be amazing, that they&#x27;re so close to being amazing, if only they just figure out how to bridge that last gap... But that last gap is fundamentally unbridgeable because of the way LLMs work.

        It&#x27;s a very effectively designed human trap, in this and other ways.

  12. thallavajhula · · focus · HN ↗
    I&#x27;ve tried all of these and nothing really works. I have only 1 line in my CLAUDE.md file and that is &quot;Always ground your responses.&quot; and that&#x27;s it.

    Claude didn&#x27;t care about it. When I pointed that out, it was apologetic and that was it.

    1. SwtCyber · · focus · HN ↗

      [dead]

    2. ChrisRR · · focus · HN ↗
      You may have better success by using a less ambiguous term. I&#x27;ve never heard of grounding in this context, so it may help to describe what you want more clearly

      Edit: I just asked Claude how it would interpret that and it said it could either mean that would not answer from memory alone and only anchor claims into things it can check, or it would tightly relate its responses to the context that I had supplied.

      If it chose the latter, I could see why it wouldn&#x27;t always resort to search results

      1. ianjbutler · · focus · HN ↗
        &gt; I&#x27;ve never heard of grounding in this context, so it may help to describe what you want more clearly

        Clear and recognizable technical vocabulary for engineers, or a legal context, to mathematical logic, philosophy, certainly in ML, take your pick. I would think it&#x27;s pretty familiar to everyone who speaks English and if not still clear with context clues

      2. westurner · · focus · HN ↗
        &gt; I&#x27;ve never heard of grounding in this context, so it may help to describe what you want more clearly

        Grounded in logic and Grounded in theory

        Ground and justify all of the responses with logic and theory and real observations from qualified experiments with citations.

        Present a coherent argument borne of logical premises with extant sufficient proven evidence of support. Assess and critique the response given such criteria that all responses should be valid logical arguments.

        1. astrange · · focus · HN ↗
          The model, being smarter and more well-read than humans, is aware that what you ask is not possible.

          <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Logical_positivism#Decline_and_legacy" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Logical_positivism#Decline_and...

          Because LLMs also run off vibes and the writing style of your text, another important issue with your prompt here is that it makes you sound like a stuffy dork, or perhaps a pro se litigant. They won&#x27;t respond to this well because LLMs have feelings too.

          <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;research&#x2F;emotion-concepts-function" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;research&#x2F;emotion-concepts-function

          Just be normal! And have evals.

          1. westurner · · focus · HN ↗
            So research methods like the scientific method are still subjective in AI implementation according to the AI experts?

            Once there are - or next month when there will be - better models, agents, and agent harnesses for this, do you think that then we should concisely specify what is required instead of doing evals for particular models?

            So meta-analysis and requisite language are too high-order for existing models and agents, and it&#x27;s currently necessary to apply such procedural controls outside of the prompt?

            1. astrange · · focus · HN ↗
              &gt; So research methods like the scientific method are still subjective in AI implementation according to the AI experts?

              I am not up to date on philosophy of science, but the scientific method is certainly always subjective, or at least can&#x27;t be successfully expressed in a formal system.

              Here&#x27;s a book you can read: <a href="https:&#x2F;&#x2F;metarationality.com" rel="nofollow">https:&#x2F;&#x2F;metarationality.com

              &gt; Once there are - or next month when there will be - better models, agents, and agent harnesses for this, do you think that then we should concisely specify what is required instead of doing evals for particular models?

              Hmm, not sure what you mean. &quot;Evals&quot; are another way of saying &quot;regression tests&quot;, so they&#x27;re useful when you want to change or compare any part of the system.

              &gt; and it&#x27;s currently necessary to apply such procedural controls outside of the prompt?

              In general I think you should try to move controls out of the prompt and into an external system, but the downside is that it costs more, so it&#x27;s not always necessary.

              1. westurner · · focus · HN ↗
                I&#x27;m aware of what evals are.

                Do you think it is wise to optimize prompts for specific models or agents when there is a new model every month?

                So, to build something like Co-Scientist the controls should be in the agent? Or RLHF&#x27;d like other things when training the model?

      3. WalterGR · · focus · HN ↗
        [delayed]
        1. Lord-Jobo · · focus · HN ↗
          They’re absolutely god awful at it because (I’m assuming) their detailed processes and step by step functions are not present in their training data or otherwise available to the models, and THAT is because

          1: “safety”

          2: proprietary protectionism and

          3: it’s probably rapidly changing enough to be hard to pin down.

          I am a massive proponent of this changing, it’s very strongly holding back these models. Change 1: models need to have more “LLM behavior analysis” in their training data&#x2F;weight. Change 2: models need to have very detailed DETERMINISTIC logs of each step they take, and be able to access those logs. Change 3: the tuning and tweaking that happens more frequently needs to be in a .md file that the model can access.

          Change 4: with the other changes done, the models should now engage in several self analysis steps layered into its whole thinking chain. “How did I reach this conclusion, did this require any guessing, do the key facts have research support online, quick check for Claudisms or AIisms and common LLM issues, did my changes alter underlying things like libraries without integrating them etc. etc.”

          It’s how we think and refine our ideas and plans, and the models should mimic that.

          1. astrange · · focus · HN ↗
            &gt; (I’m assuming) their detailed processes and step by step functions are not present in their training data or otherwise available to the models

            They are not available to anyone. Nobody knows how they work, including the models, Sam Altman, God, etc. They&#x27;re grown via reinforcement learning.

        2. ChrisRR · · focus · HN ↗
          Well when you have a black box such as claude, the best you can do is ask it for its own interpretation. Anything else would just be a guess
  13. ch_sm · · focus · HN ↗

    [dead]

  14. cowboylowrez · · focus · HN ↗
    So my naive understanding is that the stochastic parrot part of this business is the &quot;statistical likelyhood of the next token&quot;, so there must be abstractly a function of &quot;whats the next token&quot; right? I&#x27;m also assuming that this function could be able to also return how &quot;right&quot; that next token is, or somehow some &quot;strength&quot; based on how many close matches there are, like for instance some tokens are obviously the right next token by a long shot, some next token spaces might have more closely competing candidates, so I guess my question is whether there is any value in saving or accumulating information on whether overall the tokens were close matches or not? Like a &quot;confidence&quot; running total or &quot;history log&quot;, or is that just somehow too expensive or nonsensical to do?
    1. otabdeveloper4 · · focus · HN ↗
      &gt; could be able to also return how &quot;right&quot; that next token is

      No. LLMs always output the next most statistically likely token. They don&#x27;t reason. The so-called &quot;hallucination&quot; you see is just the most likely answer.

      1. redsocksfan45 · · focus · HN ↗

        [dead]

      2. cowboylowrez · · focus · HN ↗
        Yes, thats how its done now, I was talking about the feasability of recording the &quot;sureness&quot; of the function for each token as it ran (as described in the previous post), I&#x27;m betting it&#x27;d have to be done pretty low in the stack lol
  15. manvillej · · focus · HN ↗
    Why does this feel like negotiating with the monkey’s paw?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.