‹ BackHN Continuity

Thread

GPT-6 Sol and Luna

1779 points · 855 comments · OfficialTurkey

  1. simonw · · focus · HN ↗
    GPT-6 Luna being half the price of GPT-5.6 Luna is a really big deal.

    Here&#x27;s GPT-6 Luna pelicans: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F40d129fc140faca378b9c9f4f16c6ec2" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    And GPT-6 Sol: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    Scroll to the bottom for the GPT-6 Sol max one: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2#response-5" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    For comparison, here are the pelicans I got for GPT-6 Astra: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht... - I still like the Astra Max one best.

    Here&#x27;s a comparison grid showing all of the GPT-6 and GPT-5.6 pelicans at all effort levels: <a href="https:&#x2F;&#x2F;static.simonwillison.net&#x2F;static&#x2F;2026&#x2F;gpt-6-and-5.6.html" rel="nofollow">https:&#x2F;&#x2F;static.simonwillison.net&#x2F;static&#x2F;2026&#x2F;gpt-6-and-5.6.h...

    The grid is actually really interesting, because it shows that the 5.6 family default to brighter colors than the 6 family.

    1. gizmodo59 · · focus · HN ↗
      6-luna is at the pareto for most of the tasks! I dont know how they make money here but its insane value from a closed source model. I&#x27;d go further and say it makes no sense (privacy, sovereignty etc aside) to use many other models as its not only expensive but also many providers don&#x27;t have that much GPUs to serve at a significant volume. <a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings?view=month#top-models" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;rankings?view=month#top-models 5.6 luna is already the most used model this month.
      1. krat0sprakhar · · focus · HN ↗
        Can&#x27;t agree more. Between 5.6 Luna and Gemini 3.8 flash I&#x27;m so happy for the value I&#x27;m getting for my dollar (subscription pricing not API pricing) :)
        1. jadbox · · focus · HN ↗
          Gemini 3.8 Flash looks like its better than v7 Luna&#x2F;Sol on DeepSWE v1.1 while at $0.75 per million input tokens and $3.75 per million output tokens. Luna is much cheaper, but Flash has nearly Astra&#x27;s performance for under the price of Sol ($2&#x2F;$10).
          1. krat0sprakhar · · focus · HN ↗
            TBH: I really like how fast 3.8 Flash is... Once I have clear plan, I feel quite confident in delegating large parts of implementation to Flash and Luna
            1. mgkimsal · · focus · HN ↗
              Maddening for a bit - I&#x27;ve got problems that Flash is better on, and some Luna is better on, but I generally don&#x27;t know until one has wasted time&#x2F;tokens. Then I switch to the other one and... it&#x27;s often just... bam - done. Correctly. I can&#x27;t find the patterns ahead of time to determine what model I should be using first. :&#x2F; That said, I&#x27;ve been alternating between both the last month or so and they&#x27;ve both been pretty good compared to earlier models.
              1. Kostchei · · focus · HN ↗
                codex seems pretty solid on review, flash is fast on basics but makes more mistakes&#x2F;errors, Claude is just to picky for me
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.