‹ BackHN Continuity

Thread

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

1066 points · 953 comments · crorella

  1. minimaxir · · focus · HN ↗
    > Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing

    This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.

    1. TuxSH · · focus · HN ↗
      Exactly half as expensive as Opus 5.5 in every API pricing metric
      1. bigwheels · · focus · HN ↗
        And half as good. I didn't have great experiences with Anthropic models in the past, but Opus 5.5 seems to have turned a major corner. It is churning through tasks significantly more quickly and efficiently.

        Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.

        Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).

        1. dotancohen · · focus · HN ↗

            > Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.
          
          That's far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural. What is "something difficult" in your workflow?
          1. peterbell_nyc · · focus · HN ↗
            You HAVE to have a set of personal evals for each class of task you want to use models against at scale so you can test plausible candidates and compare output on your work against your evals.

            There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.

          2. Starlevel004 · · focus · HN ↗
            > OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural.

            This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.

          3. notatoad · · focus · HN ↗
            My side by side evaluation this week was to build a tool for mounting my app’s UI components in a headless chrome and feeding mock data into them, for the purpose of taking screenshots for help docs. Not super complicated, but a real task I needed done.

            I have the task to codex first, it took a couple back and forth prompts to define the project and then it worked for a bit and to took a couple more prompts before I decided it was good enough - not perfect, but close. It re-implemented some wrapper components in a simplified way that lost some of the UI, but it would work.

            Opus 5.5 took the same prompt with no back and forth, it just went off a built a tool that takes pixel-perfect screenshots of exactly what my app looks like.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.