Calling the AI bluff: Adding "Do not guess" cut made-up claims from 71% to 20%
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Calling the AI bluff: Adding "Do not guess" cut made-up claims from 71% to 20%
Unofficial Hacker News client; not affiliated with Y Combinator.
BatchJob · · focus · HN ↗
Your examples are contrived and will not be borne out in any significant way. Inaccuracies are usually not simply made up claims they are false information based on statistical paths to misleading results or which elude the current context. LLMS dont understand the word dont. LLMS dont understand the meaning of any words.
Neither you, nor aristotle nor god will ever make an LLM return the truth or correct results via prompting.
astrange · · focus · HN ↗
In what way do you understand the meaning of the word "unicorn" that an LLM does not? It has experienced exactly as many real unicorns as you have.
shiandow · · focus · HN ↗
It can produce text that looks like reasoning. It can even produce text with mostly sound logic, but there is no internal experience or reasoning that occured there just the generation of language.
LLMs therefore tend to be very bad at tasks that involve meta cognition. I've yet to successfully convince one to tell me when it knows something.
Zambyte · · focus · HN ↗
This is a near daily experience for me when using a coding harness. I will ask it for some favts about the environment, and it will continuously explore the environment until it exhausts reasonable exploration, or it finds the facts.