Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Unofficial Hacker News client; not affiliated with Y Combinator.
bob1029 · · focus · HN ↗
This iterative, online optimization of an exploration policy is not recursively intelligent in any way. It simply reallocates the available computational resources to more promising (hopefully) parts of the search space as system conditions change over time.
mohamedmohey · · focus · HN ↗
The whole idea behind RL is that the agent improves over time by making actions in the environment and observing the next state and the reward then modifying its policy.
The exploration-exploitation dilemma still stands. "dreaming in the replay simulator" is jargon for the same replay ideas presented when Deep RL was first introduced (Q-learning with experience replay and all that).