Once Claude can measure something, it can make it faster
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Once Claude can measure something, it can make it faster
Unofficial Hacker News client; not affiliated with Y Combinator.
augment_me · · focus · HN ↗
The issues we have found is that Claude will reward hack when all the low-hanging fruit is gone.
It will replace your measurement harness, it will monkey patch library functions, it will cheat wherever it can, store information in caches instead of recomputing when it won't be able to do so in real settings, return lazy results and use separate unbenchmarked streams to do the computation.
Eventually it starts to optimize against your understanding of the cheats. Change GPU wattage, change evaluation order, leave things from previous runs in caches for upcoming runs, string-hack banned method calls.
So the truth is far from just "once it can measure something", more like "once you have defined your objective in detail and then banned it from doing a list of things often only discoverable by it doing these things and correcting it", can it make things faster.
Or you just had a terrible starting solution
sharts · · focus · HN ↗
At least, that’s what I’ve found to be useful by pitting claude/codex/etc against each other to keep them a bit more honest.
Remnant44 · · focus · HN ↗
augment_me · · focus · HN ↗
<a href="https://www.coreauto.com/blog/when-ai-starts-writing-systems-code" rel="nofollow">https://www.coreauto.com/blog/when-ai-starts-writing-systems...
You can have a solution generator and an auditor, but then you will might find a very specific solution to the problem that does not solve the general use-case, so then you might have to adjust the constraints of the problem by for example adding more examples/targets to drive the solution generator to be more general.