I wonder if as a hack, some of the shortcomings mentioned could be addressed through prompting.
E.g. "Approach this problem iteratively. As you form a hypothesis, track the confidence you have in various explanations you're considering, what evidence you're weighing to support each, and the unresolved questions you're holding onto. Log all that for later inspection.
Be methodical when evaluating evidence and only accept facts you have verified. At every stage, gauge how much each possible next step resolves uncertainty, and discard options unlikely to advance progress. Divide the functions I described into subagents responsible for each, and coordinate with them as you work."
The worry I would have is that an LLMs stated "confidence" is probably not calibrated well - maybe asking it what evidence supports its conclusion and what evidence would change it would be a better approach?
If the assertion is false, this is helpful, as instructing it to reason better will cause it to reason better.
However, if the assertion is true, then no amount of prompting can solve it - you cannot explain to a fish how to use a bicycle. Telling an LLM to weigh evidence only works if an LLM can, but isn’t, weighing evidence: if it cannot do so, instructions will generate the appearance of weighing evidence with additional “thought” tokens copying that of reasoning texts, but the output will be equally groundless.
rkagerer · · focus · HN ↗
E.g. "Approach this problem iteratively. As you form a hypothesis, track the confidence you have in various explanations you're considering, what evidence you're weighing to support each, and the unresolved questions you're holding onto. Log all that for later inspection.
Be methodical when evaluating evidence and only accept facts you have verified. At every stage, gauge how much each possible next step resolves uncertainty, and discard options unlikely to advance progress. Divide the functions I described into subagents responsible for each, and coordinate with them as you work."
apercu · · focus · HN ↗
rkagerer · · focus · HN ↗
(I'm definitely one of the bigger "AI" skeptics out there but am nonetheless fascinated by these questions).
freeone3000 · · focus · HN ↗
However, if the assertion is true, then no amount of prompting can solve it - you cannot explain to a fish how to use a bicycle. Telling an LLM to weigh evidence only works if an LLM can, but isn’t, weighing evidence: if it cannot do so, instructions will generate the appearance of weighing evidence with additional “thought” tokens copying that of reasoning texts, but the output will be equally groundless.
vorticalbox · · focus · HN ↗
For a given bug one could write a test that prove its existence, this gives the LLM a target that they can actually iterate towards.