Vote on which of Hacker News' challenges for AI have been met
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Vote on which of Hacker News' challenges for AI have been met
Unofficial Hacker News client; not affiliated with Y Combinator.
bmenrigh · · focus · HN ↗
jerf · · focus · HN ↗
jerf, 2024: "If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn't what I was talking about as "high level math".
"I also am not surprised by "Consider a generation function" coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit "have you considered using wood?" is not a system that can build a house autonomously.
"It especially won't seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs."
The voting gloss: "An AI fully solves a research-level math problem on its own, not just suggesting an approach."
Yes, I'm satisfied. I don't even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from "lol, can't add two six-digit numbers" to research-math level almost overnight in comparison.
48844858 · · focus · HN ↗
CamperBob2 · · focus · HN ↗
yorwba · · focus · HN ↗
CamperBob2 · · focus · HN ↗
They are the kind of problems that, if your teacher was anything like mine, were usually skipped in order to keep the slower students from bogging down the class as a whole. 0.996 (MiMo-V2.5) is substantially better than what the vast majority of humans would do.
If you limited the question to adding arbitrary pairs of numbers of reasonable size, I imagine quite a few models could get to 1.000.