It does vey well at one shotting a PacMan clone, pretty much perfect.
<a href="https://jonclegg.github.io/pacman-bakeoff/entries/claude-sonnet-5-5.html" rel="nofollow">https://jonclegg.github.io/pacman-bakeoff/entries/claude-son...
2nd only to Opus 5.5, which is perfect.
<a href="https://jonclegg.github.io/pacman-bakeoff/entries/claude-opus-5-5-r2.html" rel="nofollow">https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results:
<a href="https://jonclegg.github.io/pacman-bakeoff/" rel="nofollow">https://jonclegg.github.io/pacman-bakeoff/
I made ~10 games with opus 5.5 (all multiplayer web games over web sockets).
About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like "the blaster weapon is way too powerful, divide it's hit points by 10" or "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS")
Also, if you then resume your conversation, and the LLM discovers the change, it might revert the "accidental" change.
If your working with an active context, and changes you do then need to be conversed back to the agent, and even then, it might still find it jarring and wrong.
I actually haven't had this happen in several months, they have gotten much better at just ignoring changes they haven't made in my experience. Unless it's directly conflicting with something they are already working on of course then they might change it but at least they will note it
Wow, that "bake off" page is better than any coding benchmark I've seen! You can really sense the strengths and weaknesses of each model/harness combo.
Cool page and benchmark idea! Would be nice if there was some kind of grading the results, maybe on different criteria (aesthetic, implementation complexity, correctness, ...). Of course as a one-shot and greenfield benchmark the results are not indicative for all kinds of usage patterns. But as some sibling said, maybe they can be indicative on some general characteristics (especially since the task is so open-ended).
Just added! I had Opus 5.5 look at them, not a perfect way to score them but it's close-ish -- Best would be a ELO, where people play both and rank a winner, but I don't know if people want to bother doing that.
Pretty cool. Using Opus as a judge should be more than good. Interesting to see more stats like the result file size too, but the grading criteria are maybe a bit skewed towards UX. Would like to see more code quality.
Very cool. I'd love to see someone with access to plenty of token$ make something similar for the "Browser Desktop OS" test. That seems like a pretty comprehensive test thats also fun to test just like this!
Well, I thought you meant the Pacman package manager for a second, still impressive. I've noticed the last few versions of Opus have produced better video game coding output.
thefourthchime · · focus · HN ↗
2nd only to Opus 5.5, which is perfect. <a href="https://jonclegg.github.io/pacman-bakeoff/entries/claude-opus-5-5-r2.html" rel="nofollow">https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: <a href="https://jonclegg.github.io/pacman-bakeoff/" rel="nofollow">https://jonclegg.github.io/pacman-bakeoff/
judge2020 · · focus · HN ↗
thefourthchime · · focus · HN ↗
thefourthchime · · focus · HN ↗
copperx · · focus · HN ↗
londons_explore · · focus · HN ↗
About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like "the blaster weapon is way too powerful, divide it's hit points by 10" or "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS")
copperx · · focus · HN ↗
igleria · · focus · HN ↗
this is literally faster to do it yourself
> "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS"
these would not.
weird-eye-issue · · focus · HN ↗
It's literally not unless you already know exactly where it is in the code
ThunderSizzle · · focus · HN ↗
If your working with an active context, and changes you do then need to be conversed back to the agent, and even then, it might still find it jarring and wrong.
weird-eye-issue · · focus · HN ↗
russellbeattie · · focus · HN ↗
thefourthchime · · focus · HN ↗
ilamont · · focus · HN ↗
sixtyj · · focus · HN ↗
Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…
lukan · · focus · HN ↗
Fable 5.1 was pretty good. Even animating it:
<a href="https://news.ycombinator.com/item?id=49526704">https://news.ycombinator.com/item?id=49526704
nicce · · focus · HN ↗
sally_glance · · focus · HN ↗
thefourthchime · · focus · HN ↗
sally_glance · · focus · HN ↗
pyaamb · · focus · HN ↗
formvoltron · · focus · HN ↗
thefourthchime · · focus · HN ↗
fakedang · · focus · HN ↗
sunaookami · · focus · HN ↗
coopykins · · focus · HN ↗
giancarlostoro · · focus · HN ↗
<a href="https://wiki.archlinux.org/title/Pacman" rel="nofollow">https://wiki.archlinux.org/title/Pacman