Agents just eat these things. After last week's HN post about Cyphral Distich, I pointed Astra and Fable at some unsolved ciphers just to see whether some joker who knew nothing about the field could get the same results, and sure enough there's plenty of low hanging fruit.
Mildly interesting anecdote: when the Cyphral Distich solution popped up a few days ago, I spent about an hour with ChatGPT trying to solve it myself without looking at the proposed solution. ChatGPT opened by saying “the solution is disputed online,” and made the dispute sound fairly convincing, which struck me as odd because things like this are usually either clearly solved or clearly not.
After I gave up (mostly because ChatGPT had given me incomplete information needed to solve it) I checked the source of the dispute. It was a site very similar to this one and someone had an AI agent working on the same problem, publishing dozens or hundreds of pages of notes. The agent found the solution page and concluded it was wrong because many of the 32 source passages supposedly didn’t contain enough text.
I dug up the PDF of the book and found the mistake - whenever a passage continued onto the next page, the agent wasn’t including that continuation. The passages weren’t actually too short.
Annoying that ChatGPT can cite sources like this without being able to properly weigh their reliability.
Verifying sources is a recursive problem - where do you stop? Humans have intuitive feel for it, but agents don’t or at least not yet (I wonder if intuition is just a secondary neural net which is currently being added to the agents as we speak).
Also as a human you are able to examine agents erroneous trajectory, real or imaginary, without contaminating your own. Agent have a problem with that - as soon as someone else’s thought is in the context it can lose track of provenance and veracity. Sometimes I think we need a bloom filter to retroactively assign “dirty” flag to invalidated or questionable token spans already in the context.
> I wonder if intuition is just a secondary neural net which is currently being added to the agents as we speak
Arguably, intuition is primary neural net, the only thing an LLM has without CoT, it just is spiky so humans only notice where it's below-average and just dismiss the rest as normal. Of course it's going to lag behind in some areas compared to others.
aaymeloglu · · focus · HN ↗
<a href="https://aaymeloglu.github.io/unsolved-ciphers/" rel="nofollow">https://aaymeloglu.github.io/unsolved-ciphers/
But I got nothing on Daniel Bordeau, who in the past week seems to have built himself a whole code breaking factory!
<a href="https://dbourdeau.github.io/cyphersolver/index.html" rel="nofollow">https://dbourdeau.github.io/cyphersolver/index.html
93po · · focus · HN ↗
After I gave up (mostly because ChatGPT had given me incomplete information needed to solve it) I checked the source of the dispute. It was a site very similar to this one and someone had an AI agent working on the same problem, publishing dozens or hundreds of pages of notes. The agent found the solution page and concluded it was wrong because many of the 32 source passages supposedly didn’t contain enough text.
I dug up the PDF of the book and found the mistake - whenever a passage continued onto the next page, the agent wasn’t including that continuation. The passages weren’t actually too short.
Annoying that ChatGPT can cite sources like this without being able to properly weigh their reliability.
DenisM · · focus · HN ↗
Verifying sources is a recursive problem - where do you stop? Humans have intuitive feel for it, but agents don’t or at least not yet (I wonder if intuition is just a secondary neural net which is currently being added to the agents as we speak).
Also as a human you are able to examine agents erroneous trajectory, real or imaginary, without contaminating your own. Agent have a problem with that - as soon as someone else’s thought is in the context it can lose track of provenance and veracity. Sometimes I think we need a bloom filter to retroactively assign “dirty” flag to invalidated or questionable token spans already in the context.
LordDragonfang · · focus · HN ↗
Arguably, intuition is primary neural net, the only thing an LLM has without CoT, it just is spiky so humans only notice where it's below-average and just dismiss the rest as normal. Of course it's going to lag behind in some areas compared to others.