‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. thefourthchime · · focus · HN ↗
    It does vey well at one shotting a PacMan clone, pretty much perfect. <a href="https:&#x2F;&#x2F;jonclegg.github.io&#x2F;pacman-bakeoff&#x2F;entries&#x2F;claude-sonnet-5-5.html" rel="nofollow">https:&#x2F;&#x2F;jonclegg.github.io&#x2F;pacman-bakeoff&#x2F;entries&#x2F;claude-son...

    2nd only to Opus 5.5, which is perfect. <a href="https:&#x2F;&#x2F;jonclegg.github.io&#x2F;pacman-bakeoff&#x2F;entries&#x2F;claude-opus-5-5-r2.html" rel="nofollow">https:&#x2F;&#x2F;jonclegg.github.io&#x2F;pacman-bakeoff&#x2F;entries&#x2F;claude-opu...

    Up until very recently, all models struggled with this.

    All results: <a href="https:&#x2F;&#x2F;jonclegg.github.io&#x2F;pacman-bakeoff&#x2F;" rel="nofollow">https:&#x2F;&#x2F;jonclegg.github.io&#x2F;pacman-bakeoff&#x2F;

    1. judge2020 · · focus · HN ↗
      Oh, it coded a Pac-Man clone. The clone was so good that I thought it was premade in some way and that Sonnet was going to play PacMan.
      1. thefourthchime · · focus · HN ↗
        Yes! The point being that up until yesterday, every model struggled with this, and now they don&#x27;t.
        1. thefourthchime · · focus · HN ↗
          Your welcome!
        2. copperx · · focus · HN ↗
          &quot;this&quot; being recreating Pacman specifically, or games?
          1. londons_explore · · focus · HN ↗
            I made ~10 games with opus 5.5 (all multiplayer web games over web sockets).

            About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like &quot;the blaster weapon is way too powerful, divide it&#x27;s hit points by 10&quot; or &quot;we need a way to reconnect a player whose network dropped mid round&quot; or &quot;the GPS doesn&#x27;t work on iOS&quot;)

            1. copperx · · focus · HN ↗
              Other models fail at oneshot creation of similar games?
            2. igleria · · focus · HN ↗
              &gt; &quot;the blaster weapon is way too powerful, divide it&#x27;s hit points by 10&quot;

              this is literally faster to do it yourself

              &gt; &quot;we need a way to reconnect a player whose network dropped mid round&quot; or &quot;the GPS doesn&#x27;t work on iOS&quot;

              these would not.

              1. weird-eye-issue · · focus · HN ↗
                &gt; this is literally faster to do it yourself

                It&#x27;s literally not unless you already know exactly where it is in the code

                1. ThunderSizzle · · focus · HN ↗
                  Also, if you then resume your conversation, and the LLM discovers the change, it might revert the &quot;accidental&quot; change.

                  If your working with an active context, and changes you do then need to be conversed back to the agent, and even then, it might still find it jarring and wrong.

                  1. weird-eye-issue · · focus · HN ↗
                    I actually haven&#x27;t had this happen in several months, they have gotten much better at just ignoring changes they haven&#x27;t made in my experience. Unless it&#x27;s directly conflicting with something they are already working on of course then they might change it but at least they will note it
    2. russellbeattie · · focus · HN ↗
      Wow, that &quot;bake off&quot; page is better than any coding benchmark I&#x27;ve seen! You can really sense the strengths and weaknesses of each model&#x2F;harness combo.
      1. thefourthchime · · focus · HN ↗
        Thanks!
    3. ilamont · · focus · HN ↗
      Thank you for doing this. It is very helpful not just for capabilities but also for costs.
    4. sixtyj · · focus · HN ↗
      I have played few of them and it seems that Opus 5.5 is the first one who really made playable PacMan clone game. On mobile as well.

      Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…

      1. lukan · · focus · HN ↗
        &quot; If we compare it with pelicans that are still not-perfect…&quot;

        Fable 5.1 was pretty good. Even animating it:

        <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49526704">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49526704

    5. nicce · · focus · HN ↗
      I was able to get similar with Qwen 3.8 27B with one shot. I think this game is too well in the training data.
    6. sally_glance · · focus · HN ↗
      Cool page and benchmark idea! Would be nice if there was some kind of grading the results, maybe on different criteria (aesthetic, implementation complexity, correctness, ...). Of course as a one-shot and greenfield benchmark the results are not indicative for all kinds of usage patterns. But as some sibling said, maybe they can be indicative on some general characteristics (especially since the task is so open-ended).
      1. thefourthchime · · focus · HN ↗
        Just added! I had Opus 5.5 look at them, not a perfect way to score them but it&#x27;s close-ish -- Best would be a ELO, where people play both and rank a winner, but I don&#x27;t know if people want to bother doing that.
        1. sally_glance · · focus · HN ↗
          Pretty cool. Using Opus as a judge should be more than good. Interesting to see more stats like the result file size too, but the grading criteria are maybe a bit skewed towards UX. Would like to see more code quality.
    7. pyaamb · · focus · HN ↗
      Very cool. I&#x27;d love to see someone with access to plenty of token$ make something similar for the &quot;Browser Desktop OS&quot; test. That seems like a pretty comprehensive test thats also fun to test just like this!
    8. formvoltron · · focus · HN ↗
      oh! How about pengo, dig dug, &amp; defender?
      1. thefourthchime · · focus · HN ↗
        I&#x27;m afraid those will be too easy. I&#x27;m not sure what the next game should be...
    9. fakedang · · focus · HN ↗
      Interesting. Sonnet 5 was horrible, and Opus 5 was unplayable, but both Sonnet 5.5 and Opus 5.5 were about as close to the real thing.
    10. sunaookami · · focus · HN ↗
      GPT models really have no taste huh.
    11. coopykins · · focus · HN ↗
      Surprising how cheap Sol 6 was.
    12. giancarlostoro · · focus · HN ↗
      Well, I thought you meant the Pacman package manager for a second, still impressive. I&#x27;ve noticed the last few versions of Opus have produced better video game coding output.

      <a href="https:&#x2F;&#x2F;wiki.archlinux.org&#x2F;title&#x2F;Pacman" rel="nofollow">https:&#x2F;&#x2F;wiki.archlinux.org&#x2F;title&#x2F;Pacman

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.