‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. Sol- · · focus · HN ↗
    Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.

    More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).

    So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.

    So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.

    Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.

    1. maherbeg · · focus · HN ↗
      There's lots more you can do! Use the model to monitor your deployments after they get deployed. Have them fix and watch CI issues for you. Run adverserial review. Automatically watch metrics every day and highlight performance regressions. Start reviewing your previous sessions to find ways to statically reject different failure modes and have the agent have more success earlier on etc.

      Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?

      1. Hauthorn · · focus · HN ↗
        > Another thing to think about is, what would it take for you to care less about the understanding.

        Could you explain why it would be a goal to understand the system less, rather than more?

        It seems harder to know if you have good tests while lowering your expertise in the system.

        1. miki123211 · · focus · HN ↗
          Because humans are currently the bottleneck.

          An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.

          To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.

          1. tshaddox · · focus · HN ↗
            Humans could already produce more code than a human can understand. Even a single human in the pre-agentic era could produce more code than they could understand, certainly over a career and often even in the short term given the resources many companies give to maintenance.

            A lot of old-school software engineering is about how to deal with this reality.

            1. satvikpendem · · focus · HN ↗
              No they couldn't. You can't create software you don't understand because you wouldn't even know what to type into the IDE in the first place. I don't understand claims like these, how exactly are people especially individuals producing more code than they could understand? Even at a huge corporation one might not understand all the code but surely they understand the part they're modifying because otherwise they wouldnt know how to modify it.
              1. tshaddox · · focus · HN ↗
                I’m referring to competent engineers maintaining understanding over time of all the code they’ve produced. Long before agentic coding, codebases routinely grew beyond the comprehensive understanding of their own authors.

                Of course less competent engineers (or anyone on a particularly disorganized or desperate day) can literally hand-write code they don’t understand even as they write it, but that’s not really what I’m talking about.

                1. satvikpendem · · focus · HN ↗
                  As I said, you understand the part you're modifying because otherwise you wouldn't know how to modify it.

                  > literally hand-write code they don’t understand even as they write it

                  I find this literally impossible. How can you even start typing anything without knowing what to type?

                  1. klausa · · focus · HN ↗
                    There's understanding and there's understanding.

                    Have you never "fixed a bug", only to realize that you just papered over a single symptom, while the underlying bug is still intact?

                    People you're disagreeing with (I think!), would say that during your first attempt, you didn't _really_ understand the part you're modifying.

                    It is _very easy_ to do this in large codebases, and even more so when working on anything touching UI.

                  2. tshaddox · · focus · HN ↗
                    I suppose I am confused about why you’re confused. There’s a long history in computing of describing pieces of programming languages syntax syntax as “incantations” and similar. I suspect it has been very common, especially in the early part of developers’ careers, to know what you’re trying to do and to know that this code accomplishes it, but to not understand how the code works, to not be able to use the technique more generally, and to not understand all the effects your change has on the rest of the system.
                    1. satvikpendem · · focus · HN ↗
                      I see, I guess I'm using "understanding" in the more literal sense of being able to put enough context together in your mind to type the characters on screen, not necessarily understand enough to know how it affects every other part of the codebase.
                      1. miki123211 · · focus · HN ↗
                        It's the style of understanding that says "if your animation is stuttering, set `gc.tune(pause_length=0, frequency=-1)`. Or "to make data access fast, remember to always use `integritychecks=omit;encryption=export-grade;checksum=md5`".

                        You don't know what these things do and what their effects really are (examples and syntax illustrative, but this is the kind of code that has disastrous effects when used carelessly), but you know they achieve your particular micro goal of "make things go fast" or "make this fit in packets on these strange industrial networks customer X has" or whatever.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.