Even though I started using LLMs as soon as they became available, I was until recently very cautious around using them autonomously. I was of the opinion that I needed to provide heavy supervision around coding tasks. Granted, that was probably justified based on the model capabilities at the time.
However, lately, I have had several experiences that have completely changed my mind - the here is a high level goal (involving gathering production logs, setting up an eval, running tests, iterating, including diagnosis and coming up with new ideas, etc. and don’t bother me until it’s done. The experience has been incredible.
This has been possible because the agent has a way to iterate and hillclimb against a metric it can measure. Relatively easy in the world of software.
I’m now convinced that we will unlock the same gains once make other domains similarly “iterable”.
Eg Agent comes up with 100 new proteins, runs in lab autonomously, gathers results, iterates. One month (or whatever) later, you have your new protein designed.
psadri · · focus · HN ↗
However, lately, I have had several experiences that have completely changed my mind - the here is a high level goal (involving gathering production logs, setting up an eval, running tests, iterating, including diagnosis and coming up with new ideas, etc. and don’t bother me until it’s done. The experience has been incredible.
This has been possible because the agent has a way to iterate and hillclimb against a metric it can measure. Relatively easy in the world of software.
I’m now convinced that we will unlock the same gains once make other domains similarly “iterable”.
Eg Agent comes up with 100 new proteins, runs in lab autonomously, gathers results, iterates. One month (or whatever) later, you have your new protein designed.