‹ BackHN Continuity

Thread

GPT-6 Sol and Luna

1779 points · 855 comments · OfficialTurkey

  1. scrlk · · focus · HN ↗
    Artificial Analysis is reporting that 6 Luna scores 2 points lower on their coding index than 5.6 Luna, but is 60% cheaper:

    > In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task.

    <a href="https:&#x2F;&#x2F;x.com&#x2F;ArtificialAnlys&#x2F;status&#x2F;2102462962758033624" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;ArtificialAnlys&#x2F;status&#x2F;2102462962758033624

    Given that they had to discontinue sales of the 20x Pro plan after the Astra release due to compute constraints, I wonder if 6 Sol &amp; Luna are smaller vs their 5.6 counterparts?

    1. scrollop · · focus · HN ↗
      Can we trust AA anymore after the last debacle a week or two ago?
      1. 6thbit · · focus · HN ↗
        wait what debacle?
        1. wyre · · focus · HN ↗
          Probably referencing how when Astra came out it was only 1 point ahead of 5.6 sol.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.