‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. Sol- · · focus · HN ↗
    Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.

    More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).

    So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.

    So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.

    Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.

    1. maherbeg · · focus · HN ↗
      There's lots more you can do! Use the model to monitor your deployments after they get deployed. Have them fix and watch CI issues for you. Run adverserial review. Automatically watch metrics every day and highlight performance regressions. Start reviewing your previous sessions to find ways to statically reject different failure modes and have the agent have more success earlier on etc.

      Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?

      1. willtemperley · · focus · HN ↗
        > Another thing to think about is, what would it take for you to care less about the understanding

        Yes please, I'd like to not understand my codebase, give up my decades of experience and have a machine do everything for me. That way I can let captialism utterly steamroller me because of my paltry token stack, in comparison to the 19 year old vibe coder who has secured a new funding round for ponzi.ai

        1. maherbeg · · focus · HN ↗
          The cat's out of the bag already. We can't undo the idea of LLMs or coding agents. If training progress stopped today, we have years and years of harness improvements to extract more performance out of today's models.

          We also have open weight models too, and ways to host those at home.

          Most people don't look at the assembler output of their C++ code (I used to write win32 programs in asm!). Most people don't look at the opcode instructions or JIT output of their ruby / python code. We're starting to work at a higher level of abstraction using LLMs. It's ok to be sad about it, but just being angry about it isn't going to change that there's a new world out there with a new skill set that's needed for honing.

          1. nananana9 · · focus · HN ↗
            And where's the results of all this higher level work?

            Where's the super awesome 100x turbocharged software that's a result of everyone here having been being a 100x turbocharged programmer for the last 6 months and a 10x supercharged programmer the past 2 years?

            I still use the same software I used 2 years ago, but a bit less reliable.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.