‹ BackHN Continuity

Thread

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

213 points · 53 comments · bananaflag

  1. ahmedhossamdev · · focus · HN ↗
    The replay simulator from history for off-policy eval is clever - avoids expensive rollouts. Curious how they prevent the policy from overfitting to already-discovered branches and going stale as the search space expands?
    1. patrykorwat · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.