‹ BackHN Continuity

Thread

Learning to solve hard problems in RL for LLMs by never giving up

119 points · 9 comments · natolambert

  1. aswegs8 · · focus · HN ↗
    Seems like persistent models like OpenAI's highly persistent internal model can become really effective over time. Those are the ones that drove most of the HF-OAI incident.
    1. numeri · · focus · HN ↗
      This training technique does not relate to how persistent a model is, at all really. They sample more parallel attempts at hard problems, to increase their chances of having at least one success to learn from.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.