‹ BackHN Continuity

Thread

Sonnet 5.5

884 points · 613 comments · D2OQZG8l5BI1S06

  1. Sol- · · focus · HN ↗
    Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.

    More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).

    So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.

    So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.

    Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.

    1. miki123211 · · focus · HN ↗
      I find that "vibe coders" (that is, people who do not know anything about programming, but nevertheless produce useful tools for themselves and others) are using a lot more tokens than we do as programmers.

      I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.

      They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).

      1. alkonaut · · focus · HN ↗
        I (a programmer) just did a pretty large task with Opus 5.5 that was a perfect fit for a big token eating task. It ate into my weekly budget in a way that made me have to use anthropics one-off "reset" they offer now.

        Long story: we have a big legacy desktop app. It uses a big legacy UI component (a grid control), which we had a license for in an old version. Fast forward 20 years, and to be able to move to a new runtime for our app, we need to update the component. Someone had bought the company making the component and now charges north of $1k per developer per year. So instead of doing this, we had just lived with the very old version.

        We had long thought of writing our own control to replace the proprietary, but it was always going to be a man-year of work we thought. But I thought I'd give it a try with AI now. I told Opus: look at our uses of that control (tens of thousands of lines of code, it has over 100 instances across our User Interface). Write a new control that would compile with the exact same app syntax. First just make a dummy implementation that throws on every call. Then start implementing. Make a test suite that can run both with our new control and the proprietary control, and test everything, every function that can be called in its public interface and every state that can be inspected from the public API. Verify that everything behaves exactly the same, and lock it in with thousands of tests. Finally, check that the control _looks_ exactly the same as the proprietary one. Render to bitmaps, figure out the rendering logic from observation, such as arithmetic for padding, font sizes, and so on. Compare pixels until it's exactly the same.

        Basically: it was a mammoth coding task, but it was so extremely well specified that an LLM could easily just do it. It's a clean-room implementation of something with no tests, but we had a test double that could provide 100% of the expected behavior. The description was extremely short. "Make a new thing that works like the old thing, and prove that it does". Opus 5.5 finished this in a number of hours. 500 source files, several thousand unit tests, and html reports with image diffs from the reimplementation and the original control. It did not use any disassembly or such "cheating". Only observation of the public API and the behavior. Do we need to deeply understand the implementation? Does the architecture matter? Not much in this case I'd argue. It was a black box to begin with and it remains a black box. If we notice a bug, we can always point it to the original proprietary control and say "there's a behavioral difference when doing X" and it will fix it, and lock it down with tests.

        As a programmer it's kind of chilling. I had recreated for a few tens of dollars something that would cost $1000 per year to buy. Obviously it's not a complete implementation only the parts of the API we use. It likely still has some bugs. We don't get support, we get to maintain it ourselves. But the rate of reverse engineering this thing "black box" was frightening. It hasn't created anything novel. But we must realize that as programmers some times we have man-years of work that just isn't novel. And in the past, we didn't do this work at all.

        I wonder if those who write and sell libraries like this will start having explicit no-reverse-engineering EULAs soon? Perhaps even explicitly mentioning AI/LLM use in analysis and reimplementation?_ Obviously the library we reimplemented was from 2005 so didn't mention AI... (It doesn't mention reverse-engineering either, luckily).

        1. Pannoniae · · focus · HN ↗
          Most products do in fact have an anti-reverse engineering clause in their EULA, to be fair, it's been a standard EULA term for a long while. It's just that no one cares anymore...
          1. alkonaut · · focus · HN ↗
            Yes, and usually in the form "You may not reverse engineer, decompile, or disassemble the SOFTWARE or any of its constituents, except and only to the extent that applicable law expressly permits". (This is the concrete example from this software). And as far as I understand, this means that so long as you stay short of decompilation - you can reimplement as much as you want.

            The law that covers this (in the EU) is EU Directive 2009/24/EC, where Article 5 is the reverse-engineering-without-decompilation.

            > The person having a right to use a copy of a computer program shall be entitled, without the authorisation of the rightholder, to observe, study or test the functioning of the program in order to determine the ideas and principles which underlie any element of the program if he does so while performing any of the acts of loading, displaying, running, transmitting or storing the program which he is entitled to do.

            This is pretty difficult to parse, but luckily there is a ruling from the European Court of Justice on this: SAS Institute Inc. v World Programming Ltd (Case C-406/10), delivered on May 2, 2012.

            SAS Institute claimed that World Programming Ltd (WPL) infringed its copyright by studying the behavior of the SAS software system and writing a competing program (the World Programming System) that emulated its exact functionality and used the same data file formats. WPL did not have access to SAS's source code and did not copy any of its literal text or internal structural design.

            CJEU:

            > "It must therefore be held that the copyright in a computer program cannot be infringed where, as in the present case, the lawful acquirer of the license did not have access to the source code of the computer program to which that license relates, but merely studied, observed and tested that program in order to reproduce its functionality in a second program".

            Which is a good find. But this is where I wonder if LLM-based reverse engineering is going to creep into either law (via lobbying) and/or EULA's, because this "observe every single state of the program for every single mutation" was simply not a viable mode of reverse engineering in the past. Or, it was at least always cheaper than just buying the software! Not so any more.

            Or alternatively, that programs stop having so many observable states, making more things public. But for libraries as in this case, the whole product IS the public API. Without a rich public API, the library can't be sold. And with it, I can observe it and copy it - because it's internal workings are "too simple" not to be deduced from the public API. In short: a UI control is a ton of hard-to-write but easy to copy boilerplate code. And selling this has been an industry, but I wonder if it will be for very long.

            1. Pannoniae · · focus · HN ↗
              "And as far as I understand, this means that so long as you stay short of decompilation - you can reimplement as much as you want."

              Yes, but my point is that.... go on github, you'll find tons of decomps. And many more done just privately too. One of the No Man's Sky devtalks start with "yeah we decompiled the terrain generation from this other game, implemented it in our prototype, it didn't work okay, here's how we've learnt from it to make something better". This was in 2016. More recently, this has been going on way more openly, even full AI-assisted decomps thrown up onto GitHub casually. It might be the letter of law or included in Terms of Service but no one cares really.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.