I feel this post. I am tired of "directing" agents when in reality it feels more like trying to herd a group of toddlers.
Sure they can mostly write better code than a toddler but this constant nudging and reminding and reiterating and stopping them from using the token budget of the whole company for a one off script. It gets tiring and I feel like I am losing brain power while doing it. Maybe it's faster but explosive diarrhea is also a faster way to produce shit.
I think the main problem is a conflict of objectives. AI companies need you to use more tokens so they are not going to improve that. They are optimizing the models to get as close as possible to the point in which people would stop using them because they are useless but without crossing that line.
Proof of this is the amount of unnecessary tasks that Claude does just as an excuse for not doing a good job doing the tasks that we ask it to do
Conflict of objectives is fine, it's not like we have a single provider of models vying for us to spend our money on. Something like collusion and price fixing across the industry would be a bit different.
BadBadJellyBean · · focus · HN ↗
Sure they can mostly write better code than a toddler but this constant nudging and reminding and reiterating and stopping them from using the token budget of the whole company for a one off script. It gets tiring and I feel like I am losing brain power while doing it. Maybe it's faster but explosive diarrhea is also a faster way to produce shit.
gonzalohm · · focus · HN ↗
Proof of this is the amount of unnecessary tasks that Claude does just as an excuse for not doing a good job doing the tasks that we ask it to do
zamadatix · · focus · HN ↗
sifar · · focus · HN ↗