‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. Sol- · · focus · HN ↗
    Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.

    More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).

    So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.

    So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.

    Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.

    1. miki123211 · · focus · HN ↗
      I find that "vibe coders" (that is, people who do not know anything about programming, but nevertheless produce useful tools for themselves and others) are using a lot more tokens than we do as programmers.

      I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.

      They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).

      1. inopinatus · · focus · HN ↗
        It's because they don't know data structures.

        "Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won’t usually need your flowcharts; they’ll be obvious." - Fred Brooks, The Mythical Man-Month (1975).

        and essentially the same sentiment, three decades later:

        "Bad programmers worry about the code. Good programmers worry about data structures and their relationships." - Linus Torvalds, git mailing list, 2006.

        These things have not changed even though everything else is topsy-turvy. As-of current writing, I have yet to see an LLM make good data structure choices; they go for something that is superficially plausible but profoundly ill-considered (or rather, not considered at all), and then commonly burn tokens treating this implementation detail as a design invariant and trying to deal with the consequences by writing more code, instead of iterating directly upon the ill-fitting data at the root its problems.

        If you're wondering, "does he mean the schema of let's say a db or other persistent store, or does he mean abstract/algebraic structures", the answer is yes to both, I think coding models are today shockingly weak when it comes to design reasoning in both domains.

        Fortunately, their suggestibility means the same models will readily accept direction on the matter (perhaps even more so than on the structure of code), so I recommend doing just that, and (bonus!) this means your CS degree is still relevant.

        1. avmich · · focus · HN ↗
          Watch LLM start paying attention to data structures.
          1. inopinatus · · focus · HN ↗
            A coding model that can make genuinely well-considered data structure choices won't be a LLM, it'll be something more general.
            1. hathawsh · · focus · HN ↗
              While I agree that a coding model (such as Opus) by itself tends to act very shallowly, when it's driven by a harness like Claude Code, the combination seems to be a far more general thing than a LLM. It's capable of consistently making excellent data structure and architectural choices over large code bases. It imitates thinking about anything and it can drive itself for hours.

              Honestly, if I simply fed it a sense of presence (I would repeatedly tell it what's going on right now and ask it to react if it thinks it should), it would feel eerily like AGI.

            2. logicchains · · focus · HN ↗
              Coding models can already make well-considered data structure choices if given all the relevant context, but a non-programmer doesn't know the context to give it.
          2. stymaar · · focus · HN ↗
            It's not going to happen naturally, the labs first need to implement a reinforcement learning pipeline that promotes it.
            1. potbelly83 · · focus · HN ↗
              Falling back on a RL pipeline to cover gaps always strikes me as a more sophisticated version of the mechanical turk. If what we had was truly AGI wouldn't they be able to derive this from the data they already have.
              1. sdeframond · · focus · HN ↗
                Why would we care wether something truly is AGI or not?

                It is useful. It may be dangerous. It has an impact. I care about that.

          3. jappgar · · focus · HN ↗
            If you frame the conversation in those terms, they will.

            One of the problems is that by default, they'll avoid changing data structures or architecture that is already written down.

            Like a junior dev, they're correctly cautious about breaking things, so they prefer to write more code instead.

            1. sdeframond · · focus · HN ↗
              > If you frame the conversation in those terms, they will.

              Indeed I realized recently that, when we complain about LLMs producing slop, that's in part because we dont ask them to refactor.

              Coding agents won't, on their own, make a big change the user did not ask for. And this is fine.

          4. [deleted] · · focus · HN ↗

            [deleted]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.