‹ BackHN Continuity

Thread

Calling the AI bluff: Adding "Do not guess" cut made-up claims from 71% to 20%

96 points · 42 comments · FKJ

  1. thallavajhula · · focus · HN ↗
    I've tried all of these and nothing really works. I have only 1 line in my CLAUDE.md file and that is "Always ground your responses." and that's it.

    Claude didn't care about it. When I pointed that out, it was apologetic and that was it.

    1. ChrisRR · · focus · HN ↗
      You may have better success by using a less ambiguous term. I've never heard of grounding in this context, so it may help to describe what you want more clearly

      Edit: I just asked Claude how it would interpret that and it said it could either mean that would not answer from memory alone and only anchor claims into things it can check, or it would tightly relate its responses to the context that I had supplied.

      If it chose the latter, I could see why it wouldn't always resort to search results

      1. WalterGR · · focus · HN ↗
        > I just asked Claude

        How well are LLMs able to reason about their own behavior?

        Put another way, this entire post is about getting AI to not bluff. How do you know that it’s not bluffing in its response to you?

        1. Lord-Jobo · · focus · HN ↗
          They’re absolutely god awful at it because (I’m assuming) their detailed processes and step by step functions are not present in their training data or otherwise available to the models, and THAT is because

          1: “safety”

          2: proprietary protectionism and

          3: it’s probably rapidly changing enough to be hard to pin down.

          I am a massive proponent of this changing, it’s very strongly holding back these models. Change 1: models need to have more “LLM behavior analysis” in their training data/weight. Change 2: models need to have very detailed DETERMINISTIC logs of each step they take, and be able to access those logs. Change 3: the tuning and tweaking that happens more frequently needs to be in a .md file that the model can access.

          Change 4: with the other changes done, the models should now engage in several self analysis steps layered into its whole thinking chain. “How did I reach this conclusion, did this require any guessing, do the key facts have research support online, quick check for Claudisms or AIisms and common LLM issues, did my changes alter underlying things like libraries without integrating them etc. etc.”

          It’s how we think and refine our ideas and plans, and the models should mimic that.

          1. astrange · · focus · HN ↗
            > (I’m assuming) their detailed processes and step by step functions are not present in their training data or otherwise available to the models

            They are not available to anyone. Nobody knows how they work, including the models, Sam Altman, God, etc. They're emergent from the training process.

            > Change 4: with the other changes done, the models should now engage in several self analysis steps layered into its whole thinking chain. “How did I reach this conclusion, did this require any guessing, do the key facts have research support online, quick check for Claudisms or AIisms and common LLM issues, did my changes alter underlying things like libraries without integrating them etc. etc.”

            Remember inference costs per token. Do you want to pay for this every time?

        2. ChrisRR · · focus · HN ↗
          Well when you have a black box such as claude, the best you can do is ask it for its own interpretation. Anything else would just be a guess
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.