‹ BackHN Continuity

Thread

Writing Rust code that's fast by asking agents to make the code faster

114 points · 63 comments · mooreds

  1. hombre_fatal · · focus · HN ↗
    If it can be measured, then LLMs can optimize it.

    Once I had repo commands that could dump `sample` results and a cpu profiler/trace and then a benchmark tool that let me A/A + ABBA/BAAB-test the current modified git workspace against HEAD or any commit, the LLMs could just do their thing.

    And that's how my homemade terminal uses much less memory than ghostty/kitty/iterm yet has more throughput.

    AI is going to increasingly unmask people and companies who don't care about correct and performant software now that it's become so trivial to guarantee both. It used to at least be expensive and time-consuming and expertise-demanding to do those things.

    1. Capricorn2481 · · focus · HN ↗
      > If it can be measured, then LLMs can optimize it

      Then they can start attempting to optimize it. They can also spin round and round making the numbers worse because they don't actually know what to do.

      1. hombre_fatal · · focus · HN ↗
        Yeah, but that's just the scientific process of hypothesis -> evidence -> conclusion.

        You need a measurement that can falsify hypotheses and reject branches that won't work.

        Also, if all you have left in your project are performance issues that are hard to identify without flailing around (even with Fable/Astra) despite sampler/profiler reports, then you're doing really well and I wouldn't assume you're going to fare much better than the sota models in terms of stabs in the dark.

      2. minimaxir · · focus · HN ↗
        The point of this post is that this is explicitly not the case. If the metric is measured, the agent finds a way eventually (around 5 total tries typically unless it gets stuck), and learns from iterations where changes caused a regression after a revert.

        In one case I used a made-up metric (since I didn't know the exact name or if it existed) and it somehow optimized that too.

        1. Capricorn2481 · · focus · HN ↗
          I was actually emphasizing a methodology like yours. There is a lot more going on in your post than "the metric is measured." And even then, there's no guarantee the model will do anything except burn all your tokens trying stuff, it just won't commit the failed attempts.
          1. minimaxir · · focus · HN ↗
            "Attempting" implies a high risk of failure. That is not the case with the workflow described in the article, and from other comments, the approach is generally reliable as well.

            It's also not terrible on token usage for smaller projects.

            1. Capricorn2481 · · focus · HN ↗
              I don't know why I have to say this again, but I am telling you that the workflow you underlined in the article is good. But the more you argue about it, the more it sounds like you want people to take a more generalized conclusion than what you actually demonstrated.

              > "Attempting" implies a high risk of failure.

              Of course it does? Your safeguards also imply a high risk of failure. You have restrictions that just rollback everything the LLM "attempts" to do.

              That is not to say that the overall workflow is failure prone, but obviously you have setup an apparatus that allows the LLM to just shotgun attempts, whether it understands it or not. And sometimes it's not going to be able to find any solution. So it's not really appropriate for people to leave with the impression that anything measurable can be successfully optimized with LLMs.

      3. Conscat · · focus · HN ↗
        I had this experience at work trying to optimize a little high level Pytorch. It can't really get better than it already was, but LLMs were quite willing to pretend they will. The real solution is I need to open a PR for one of Pytorch's tracking issues.
      4. loeg · · focus · HN ↗
        > They can also spin round and round making the numbers worse

        "Claude, if this idea doesn't measure as an improvement (use X benchmark and a T-test), discard it and try the next idea."

        1. Capricorn2481 · · focus · HN ↗
          Yes, that is the spinning round and round part.
          1. loeg · · focus · HN ↗
            Crucially, it is easy to avoid making things worse, which was the entire thrust of my earlier comment and the part you've ignored in your response.
            1. Capricorn2481 · · focus · HN ↗
              I didn't ignore it, I just think spinning round and round is also bad.
              1. loeg · · focus · HN ↗
                Ok. I think the goalposts have shifted from your original remarks.
                1. Capricorn2481 · · focus · HN ↗
                  I don't think so? I think spending a bunch of money on agent tokens with no outcome is worth flagging, which was the substance of my comment.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.