‹ BackHN Continuity

Thread

I stopped reviewing my agents' code. Here's what I do instead

12 points · 12 comments · alexeyindeev

  1. moltar · · focus · HN ↗
    Maybe it depends on what you are working on. I work on platforms and engineering tooling. I cannot say the agent can be trusted yet in my experience. And I’d say I’m pretty advanced agent user compared to my peers. The majority of the work I do I have to steer and correct a lot.

    For example, recently I’ve been adding skills for our team. I added evals. If you ask an agent to make evals it’ll always try to bias and overfit just to make tests pass. No matter if I have it stated in every possible place not to do it and Codex instructed to catch these cases. It doesn’t work. Lots of sloppy evals get written and skills get extremely specific instructions to pass specific tests.

    I think if you are implementing trivial features in a green field product it can work to some degree for some time. But eventually it’s going to deteriorate into a mess. Death by 1000 paper cuts.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.