‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. BrenBarn · · focus · HN ↗
    A lot of the tests have the form of a goal ("the AI can do X") with a qualifier ("and the AI does not do Y"). I think a lot of the worrisome aspects of LLMs concern these latter qualifiers. We are used to human failure modes and those failure modes are to a large extent embedded in the complex system of physical reality, evolution, etc., giving them a certain stability. The LLM failure modes are often much more surprising. I'd like to see goalposts along the lines "people ask an LLM to do X 1000 times over a period of three years and it never does something bizarrely catastrophic". It's no good having it solve Navier-Stokes as long as it might also do the Huggingface breakout thing.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.