Calling the AI bluff: Adding "Do not guess" cut made-up claims from 71% to 20%
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Calling the AI bluff: Adding "Do not guess" cut made-up claims from 71% to 20%
Unofficial Hacker News client; not affiliated with Y Combinator.
thallavajhula · · focus · HN ↗
Claude didn't care about it. When I pointed that out, it was apologetic and that was it.
ChrisRR · · focus · HN ↗
Edit: I just asked Claude how it would interpret that and it said it could either mean that would not answer from memory alone and only anchor claims into things it can check, or it would tightly relate its responses to the context that I had supplied.
If it chose the latter, I could see why it wouldn't always resort to search results
westurner · · focus · HN ↗
An eval of this is likely worthwhile;
Re: "Grounded in logic" and "Grounded in theory"
Ground and justify all of the responses with logic and theory and real observations from qualified experiments with citations.
Present a coherent argument borne of logical premises with extant sufficient proven evidence of support. Assess and critique the response given such criteria that all responses should be valid logical arguments, and revise before responding
astrange · · focus · HN ↗
<a href="https://en.wikipedia.org/wiki/Logical_positivism#Decline_and_legacy" rel="nofollow">https://en.wikipedia.org/wiki/Logical_positivism#Decline_and...
Because LLMs also run off vibes and the writing style of your text, another important issue with your prompt here is that it makes you sound like a stuffy dork, or perhaps a pro se litigant. They won't respond to this well because LLMs have feelings too.
<a href="https://www.anthropic.com/research/emotion-concepts-function" rel="nofollow">https://www.anthropic.com/research/emotion-concepts-function
Just be normal! And have evals.
westurner · · focus · HN ↗
Once there are - or next month when there will be - better models, agents, and agent harnesses for this, do you think that then we should concisely specify what is required instead of doing evals for particular models?
So meta-analysis and requisite language are too high-order for existing models and agents, and it's currently necessary to apply such procedural controls outside of the prompt?
astrange · · focus · HN ↗
I am not up to date on philosophy of science, but the scientific method is certainly always subjective, or at least can't be successfully expressed in a formal system.
Here's a book you can read: <a href="https://metarationality.com" rel="nofollow">https://metarationality.com
> Once there are - or next month when there will be - better models, agents, and agent harnesses for this, do you think that then we should concisely specify what is required instead of doing evals for particular models?
Hmm, not sure what you mean. "Evals" are another way of saying "regression tests", so they're useful when you want to change or compare any part of the system.
> and it's currently necessary to apply such procedural controls outside of the prompt?
In general I think you should try to move controls out of the prompt and into an external system, but the downside is that it costs more, so it's not always necessary.
westurner · · focus · HN ↗
Do you think it is wise to optimize prompts for specific models or agents when there is a new model every month?
So, to build something like Co-Scientist the controls should be in the agent? Or RLHF'd like other things when training the model?