Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia
Unofficial Hacker News client; not affiliated with Y Combinator.
criemen · · focus · HN ↗
The highest scoring submission that won the contest had a high score of around 137k. Last week, I had GPT-6 Astra, Sol and Luna implement and hill-climb on this task, as I wanted to see how big the difference in smartness is. Luna implemented something, but never exceeded ca. 20k points, with a large variance. Sol got something in the area of the humans implementation.
Astra, which finished fastest, had a highscore of around 1.7Mm when the game seemed to fairly reliably crash. On the way, it disassembled parts of the ROM to extract information about the game.
I didn't do a ton work to document and measure the specifics, but it was very impressive.
looperhacks · · focus · HN ↗
criemen · · focus · HN ↗
Comparing strategies, though, Astra does something pretty different from the highest scoring submission. It doesn't predict the RNG, instead it reacts to the visuals on-screen, planning ahead by estimating velocity of all objects on screen. Unlike the winning solution, it actively flies the ship, whereas the winning solution basically only rotates and teleports. So Astra behaves more like a regular "perfect" player would, rather than one that breaks the PRNG.
suddenlybananas · · focus · HN ↗
criemen · · focus · HN ↗
To be fair, though, I mainly ran this experiment to distinguish models, not to see if they were better than humans or not, so for that purpose it didn't really matter if the basic conditions were the same or not.
NitpickLawyer · · focus · HN ↗
You should check the logs, there's probably a bunch of retro gaming forums that got hacked behind the scenes :D