‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. Sol- · · focus · HN ↗
    Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.

    More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).

    So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.

    So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.

    Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.

    1. miki123211 · · focus · HN ↗
      I find that "vibe coders" (that is, people who do not know anything about programming, but nevertheless produce useful tools for themselves and others) are using a lot more tokens than we do as programmers.

      I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.

      They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).

      1. gchamonlive · · focus · HN ↗
        I think it's not only a matter of token efficiency. If you don't know what you are doing development will eventually crawl to a halt invariably.

        It's the compound counter-probability of success, so even a 99% efficient model will in time accumulate so much error that without conscious cleanup and steering, it becomes really unlikely really fast that anything could be changed in the code without affecting something else, no matter how many tokens you throw at it. It's the collapse of a complex system under the weight of sheer uncertainty of what the system actually does.

        1. xbmcuser · · focus · HN ↗
          The llm are improving though maybe a year from now they can use it to fix the code
          1. gchamonlive · · focus · HN ↗
            Maybe, it'll be exciting to see, and I'm all about accessibility, but in this case I also don't think it's about model capability or intelligence, it's the low information to noise ratio in the codebase. There just won't be enough information in the code itself to know what to fix. Fix how? What should it do? I'm really not sure you can reconstruct intention from a codebase created unsupervised.
            1. noisy_boy · · focus · HN ↗
              1. Implement feature and write tests for the code

              2. Make sure tests pass

              3. <Every now and then> Review code for quality and fix - make sure tests pass.

              4. Go to #1

              Overly simplistic? Yes. But I would wager that this can go a long way, even for vibe coders.

              1. gbalduzzi · · focus · HN ↗
                I keep seeing this but I'm not sure it is effective in the long run *if unsupervised*.

                "Review quality and fix" doesn't mean a lot without context.

                Does it mean to remove unused features and simplify the underlaying code? Does it mean changing the data structures to better support future development? Does it mean improving performance because of bottlenecks?

                You are supposed to tell an LLM what your codebase needs, but if you just vibe code without knowing the code, "review quality and fix" will have unexpected results

                1. gchamonlive · · focus · HN ↗
                  [delayed]
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.