Writing Rust code that's fast by asking agents to make the code faster
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Writing Rust code that's fast by asking agents to make the code faster
Unofficial Hacker News client; not affiliated with Y Combinator.
hombre_fatal · · focus · HN ↗
Once I had repo commands that could dump `sample` results and a cpu profiler/trace and then a benchmark tool that let me A/A + ABBA/BAAB-test the current modified git workspace against HEAD or any commit, the LLMs could just do their thing.
And that's how my homemade terminal uses much less memory than ghostty/kitty/iterm yet has more throughput.
AI is going to increasingly unmask people and companies who don't care about correct and performant software now that it's become so trivial to guarantee both. It used to at least be expensive and time-consuming and expertise-demanding to do those things.
Capricorn2481 · · focus · HN ↗
Then they can start attempting to optimize it. They can also spin round and round making the numbers worse because they don't actually know what to do.
hombre_fatal · · focus · HN ↗
You need a measurement that can falsify hypotheses and reject branches that won't work.
Also, if all you have left in your project are performance issues that are hard to identify without flailing around (even with Fable/Astra) despite sampler/profiler reports, then you're doing really well and I wouldn't assume you're going to fare much better than the sota models in terms of stabs in the dark.
minimaxir · · focus · HN ↗
In one case I used a made-up metric (since I didn't know the exact name or if it existed) and it somehow optimized that too.
Capricorn2481 · · focus · HN ↗
minimaxir · · focus · HN ↗
It's also not terrible on token usage for smaller projects.
Capricorn2481 · · focus · HN ↗
> "Attempting" implies a high risk of failure.
Of course it does? Your safeguards also imply a high risk of failure. You have restrictions that just rollback everything the LLM "attempts" to do.
That is not to say that the overall workflow is failure prone, but obviously you have setup an apparatus that allows the LLM to just shotgun attempts, whether it understands it or not. And sometimes it's not going to be able to find any solution. So it's not really appropriate for people to leave with the impression that anything measurable can be successfully optimized with LLMs.
Conscat · · focus · HN ↗
loeg · · focus · HN ↗
"Claude, if this idea doesn't measure as an improvement (use X benchmark and a T-test), discard it and try the next idea."
Capricorn2481 · · focus · HN ↗
loeg · · focus · HN ↗
Capricorn2481 · · focus · HN ↗
loeg · · focus · HN ↗
Capricorn2481 · · focus · HN ↗