‹ BackHN Continuity

Thread

GPT-6 Sol and Luna

1779 points · 855 comments · OfficialTurkey

  1. simonw · · focus · HN ↗
    GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

    Here&#x27;s GPT-6 Luna pelicans: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F40d129fc140faca378b9c9f4f16c6ec2" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    And GPT-6 Sol: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    Scroll to the bottom for the GPT-6 Sol max one: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2#response-5" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    For comparison, here are the pelicans I got for GPT-6 Astra: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht... - I still like the Astra Max one best.

    Here&#x27;s a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: <a href="https:&#x2F;&#x2F;static.simonwillison.net&#x2F;static&#x2F;2026&#x2F;gpt-6-and-5.6.html" rel="nofollow">https:&#x2F;&#x2F;static.simonwillison.net&#x2F;static&#x2F;2026&#x2F;gpt-6-and-5.6.h...

    The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.

    1. gtirloni · · focus · HN ↗
      What&#x27;s the relevance of the pelican benchmark when models probably saw it during training? Didn&#x27;t OpenAI stop testing against SWE-Something because it was tainted?
      1. genidoi · · focus · HN ↗
        It&#x27;s not a benchmark, it is a meme benchmark.
        1. a3w · · focus · HN ↗
          Memes are arguably the web scale of benchmarks.
          1. ljm · · focus · HN ↗
            AI reproducing Xtranormal video clips like NodeJS Is Web Scale should be the new benchmark.

            If the dialogue is slop and not like the old memes then it fails.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.