‹ BackHN Continuity

Thread

ImpactGate: A merge gate that scores the structural decay AI adds

37 points · 48 comments · sagenschneider

  1. MichaelNolan · · focus · HN ↗
    Maybe I missed it, but it look like this has just a single metric. Maybe instead of making a new project, you could try to get this metric added to a existing tool like <a href="https:&#x2F;&#x2F;dekobon.github.io&#x2F;big-code-analysis&#x2F;index.html" rel="nofollow">https:&#x2F;&#x2F;dekobon.github.io&#x2F;big-code-analysis&#x2F;index.html which already has dozens of metrics.
    1. stingraycharles · · focus · HN ↗
      That project in itself looks very interesting. How are people using it, any examples of how people get this into an actual report &#x2F; CI test &#x2F; benchmark &#x2F; whatever ?
      1. MichaelNolan · · focus · HN ↗
        Code metrics in general aren’t that widely used. I’ve only ever worked at one place (a bank) that tracked it, and that was only because sonarcube had it built in.

        While a lot of metrics make intuitive sense, we don’t have that much hard evidence to prove or disprove their value. Part of it is the whole “if a metric becomes a target, it ceases to be a good metric” thing. Adding the checks to a large existing project probably has negative value. But I think it’s worth doing for greenfield projects.

        For humans, these should just be advisory. But for LLMs I’m happy enough to make it a blocking check.

        I keep thinking of doing an experiment where I give the same LLM the same problem, and only change which metric is enforced. And then see if any of them have a noticeable effect on correctness&#x2F;maintainability.

        &gt; any examples of how people get this into an actual report &#x2F; CI test &#x2F; benchmark &#x2F; whatever ?

        Yeah they have examples of adding it to CI, or local checks, generate html reports, etc in their docs.

        1. sagenschneider · · focus · HN ↗
          Yep, this all actually started because of experimenting with my own open source project <a href="https:&#x2F;&#x2F;officefloor.net" rel="nofollow">https:&#x2F;&#x2F;officefloor.net (giving full disclosure)

          I was testing the additive pipeline style of OfficeFloor against the mutative handler style of Spring. I was looking to see what factors could be used to allow AI to make long on going changes (experiment is 60 changes to an end point, where all add functionality and every 4th change is mutative on existing rules). Then I watch how AI manages to make the 60 changes in each architecture.

          I&#x27;ve done many runs and you are quite right about Goodhart effect in giving it the metric. Never knew Spring code could be written so badly.

          I&#x27;ve tried runs with better prompting also and I&#x27;m starting to find the key factor is actually the architecture itself.

          From my initial findings, it&#x27;s seeming that additive pipeline architectures hold up much better against AI slop than our typically single method web handler architectures.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.