Vote on which of Hacker News' challenges for AI have been met
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Vote on which of Hacker News' challenges for AI have been met
Unofficial Hacker News client; not affiliated with Y Combinator.
hatthew · · focus · HN ↗
Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".
intelkishan · · focus · HN ↗
hatthew · · focus · HN ↗
I didn't really design a comprehensive test, just listed a few examples of the sort of thing I imagine when I head the name "humanity's last exam". So take it with a grain of salt.
andai · · focus · HN ↗
I think non-RLHF'd LLMs (i.e. pretrained text completion models) sound natural enough to pass the Turing test, but I don't know if anyone has tested them for that. (Also I'm not sure how to come by base models without post-training crap, even the "base" models of recent releases start spamming assistant-type text constantly, i.e. they're clearly putting it in the pretraining data.)
If I'm right on that then we hit that benchmark like five years ago.
The egg thing, probably 2030-ish.
Kim_Bruning · · focus · HN ↗
The Claude answer is 'neutral', which is sure to anger people at either extreme of the AI debate (and does).
lostmsu · · focus · HN ↗
No, they will entirely forget things you ask them to remember in the beginning of a conversation. That requires at least an ability to compact context.