‹ BackHN Continuity

Thread

I stopped reviewing my agents' code. Here's what I do instead

12 points · 12 comments · alexeyindeev

  1. bryzaguy · · focus · HN ↗
    I truly empathize with this idea. It does feel like things are improving. I was scanning the code generated and felt the same way. Like, woah, maybe I'm in the clear and I can hand the reigns over. Do I even need to look at the code anymore?

    Then I ran into a few bugs and, as I dug in, what initially looked like reasonable code suddenly seemed strange when looking closer. After some help from the agent to understand the intent, it became clear the design was poor, explained a few failures, and resulted in an order of magnitude more (reasonable looking) code to compensate which now I have to sift through.

    To make matters worse, I could not get the agent to divorce itself from its wrong decisions, even with my explicit instruction detailing how the code should read. It would keep rewriting the same bad ideas then go into other distracting tangents which I'd have to correct. It felt just like how conversations in ChatGPT that would get stuck once it was in the context making me imagine we're still driving the same car but with a new paint job.

    I finally wrote the code myself. It was more helpful after that. Hopefully, I'm able to resist the temptation and stay vigilant in my reviews. For context, I am using Codex Astra XHigh but maybe Opus 5.5 really is better?

    1. aytigra · · focus · HN ↗
      One necessary thing you can do with claude (probably gpt has something similar) is to run /code-review on a design-plan document and iterate until it has no serious gaps (which in most cases I have to personally think through to make reasonable decisions and prevent llm from inventing requirements), then code generation is pretty smooth and following reviews are mostly handling various edge cases.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.