Writing Rust code that's fast by asking agents to make the code faster
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Writing Rust code that's fast by asking agents to make the code faster
Unofficial Hacker News client; not affiliated with Y Combinator.
hombre_fatal · · focus · HN ↗
Once I had repo commands that could dump `sample` results and a cpu profiler/trace and then a benchmark tool that let me A/A + ABBA/BAAB-test the current modified git workspace against HEAD or any commit, the LLMs could just do their thing.
And that's how my homemade terminal uses much less memory than ghostty/kitty/iterm yet has more throughput.
AI is going to increasingly unmask people and companies who don't care about correct and performant software now that it's become so trivial to guarantee both. It used to at least be expensive and time-consuming and expertise-demanding to do those things.
Capricorn2481 · · focus · HN ↗
Then they can start attempting to optimize it. They can also spin round and round making the numbers worse because they don't actually know what to do.
hombre_fatal · · focus · HN ↗
You need a measurement that can falsify hypotheses and reject branches that won't work.
Also, if all you have left in your project are performance issues that are hard to identify without flailing around (even with Fable/Astra) despite sampler/profiler reports, then you're doing really well and I wouldn't assume you're going to fare much better than the sota models in terms of stabs in the dark.
minimaxir · · focus · HN ↗
In one case I used a made-up metric (since I didn't know the exact name or if it existed) and it somehow optimized that too.
Capricorn2481 · · focus · HN ↗
minimaxir · · focus · HN ↗
It's also not terrible on token usage for smaller projects.
Capricorn2481 · · focus · HN ↗
> "Attempting" implies a high risk of failure.
Of course it does? Your safeguards also imply a high risk of failure. You have restrictions that just rollback everything the LLM "attempts" to do.
That is not to say that the overall workflow is failure prone, but obviously you have setup an apparatus that allows the LLM to just shotgun attempts, whether it understands it or not. And sometimes it's not going to be able to find any solution. So it's not really appropriate for people to leave with the impression that anything measurable can be successfully optimized with LLMs.
Conscat · · focus · HN ↗
loeg · · focus · HN ↗
"Claude, if this idea doesn't measure as an improvement (use X benchmark and a T-test), discard it and try the next idea."
Capricorn2481 · · focus · HN ↗
loeg · · focus · HN ↗
Capricorn2481 · · focus · HN ↗
loeg · · focus · HN ↗
Capricorn2481 · · focus · HN ↗
dasil003 · · focus · HN ↗
This isn't a new problem by any means, but now that code is cheap, it means instead of getting frustrated with engineering and their pesky unimportant details, people will get frustrated with the AI and it's pesky unimportant details.
hombre_fatal · · focus · HN ↗
I think it's one reason why ADRs are an important of a software project, especially with LLMs. You need a place were you can document invariants, why you have them + the rejected ideas and acceptable risks.
It helps smart agents like Fable help you decide on trade-offs and it's kind of incredible to witness that happening.
ashkankiani · · focus · HN ↗
This is why I'm not worried about being replaced for now or the forseeable future. For all of the improvements they've made, this part just never seems to change. They could slap another heuristic prompt for the edge case, but eventually it'll revert to the mean again.
I think there is a way to use LLMs to help with programming, but not when I'm not the driver in the seat writing the tests and deciding the architecture. Also I would never ship code written by them as the final product for anything I care about. Since I, like most people, find reading code to be arduous. The more fun thing to do is to force yourself to rewrite it all, treating the LLM's work as a rough draft.
loeg · · focus · HN ↗
They can, in fact, generate plausible performance optimization ideas on their own.
hombre_fatal · · focus · HN ↗
Make sure you process doesn't depend on anyone reading your mind.
When I run into things like this, it becomes a one-liner in my instructions/harness or in the canned prompt/skill I use that sets off a process.
In this case, I instruct agents to proactively build/improve diagnostic tooling if it would help them with their task + if it meets a bar of generalization/reusability (else it should be an ephemeral probe that gets abandoned at the end of the solution).