‹ BackHN Continuity

Thread

ImpactGate: A merge gate that scores the structural decay AI adds

37 points · 48 comments · sagenschneider

  1. MichaelNolan · · focus · HN ↗
    Maybe I missed it, but it look like this has just a single metric. Maybe instead of making a new project, you could try to get this metric added to a existing tool like <a href="https:&#x2F;&#x2F;dekobon.github.io&#x2F;big-code-analysis&#x2F;index.html" rel="nofollow">https:&#x2F;&#x2F;dekobon.github.io&#x2F;big-code-analysis&#x2F;index.html which already has dozens of metrics.
    1. sagenschneider · · focus · HN ↗
      Yes, I&#x27;m doing my own research on AI augmented pipelines <a href="https:&#x2F;&#x2F;blog.officefloor.net" rel="nofollow">https:&#x2F;&#x2F;blog.officefloor.net . I actually found most code quality tools look for bugs and complexity, but nothing much about cohesive erosion. The nice thing about this metric, is that it determine the files where the erosion is occurring. I turned it into a GitHub action to make it easier to access to get wider feedback on the metric. The GitHub action triggers on your merge request and tells you the files where erosion is occurring to refactor. This stops erosion before it gets too expensive to change (big refactors or rewrite). Yes, happy to work with others to get the metric into other tools.
      1. catlifeonmars · · focus · HN ↗
        What exactly is “cohesive erosion”?
        1. sagenschneider · · focus · HN ↗
          Comes from the basic Computer Science principals of High Cohesion and Low Coupling.

          High cohesion means the functionality of a component are closely related and focused on performing a single well defined task. Basically single classes for single purposes.

          Erosion of this is when classes start doing to many things, in the case of God classes.

          The Change Impact formula looks at a way of detecting when the cohesion is eroding and flagging it on a change (as the pull&#x2F;merge request itself should generally be single focus cohesive change)

      2. apercu · · focus · HN ↗
        I&#x27;ve never encountered that term before (cohesive erosion) but I like it, if I&#x27;m interpreting it correctly.

        Do you mean like the hyper focus an LLM puts on the task in front of it so you end up with drift (duplicated concepts&#x2F;multiple ways of doing things, terminology drift (e.g., now we have &quot;customer&quot; and &quot;client&quot;). That sort of thing?

        1. sagenschneider · · focus · HN ↗
          When you think about a god class or god method, it occurs over time by adding more than a single responsibility.

          Yes, there are generally complex algorithms but they usually are not things developers write (imported from libraries).

          What is usually going on in the god class&#x2F;method is that things keep getting added to it. These things should be separated out. So the cohesiveness of the class&#x2F;method erodes into doing too many things.

          The idea of the Change Impact formula is to catch this early so you start refactoring to separate out into classes with single cohesive purposes.

          The problem with AI is it handles complexity really well and will happily keep piling changes into god classes&#x2F;methods reaching ridiculous CC levels (have see over 200). Previously developers would get annoyed and do the refactor. But with AI these days, changes are happening faster. So Change Impact is to try to monitor the cohesive erosion.

      3. visarga · · focus · HN ↗
        I just dump all user messages from all sessions in a project into a flat .md file and have agents synthesize the user&#x27;s intent. Then, using that extracted intent, the agents review code and tests. I call this a retro&#x2F;reflection pass. It checks whether the code matches the intent and whether the tests match the code.

        Compactly formatted user messages are something an agent can ingest in a few minutes, even if they are thousands of lines long. And the quality of those messages is great: they don&#x27;t track what the agent does well, only what changes and what breaks.

        Having this top-down view helps a lot. Usually, within a session and deep into a task, the agent loses the global perspective and optimizes for local success. I find it weird there is no harness that treats user messages as high value signal (except my own, of course, I have it, <a href="https:&#x2F;&#x2F;github.com&#x2F;horiacristescu&#x2F;playbook-harness" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;horiacristescu&#x2F;playbook-harness).

        1. sagenschneider · · focus · HN ↗
          Keeping all the specifications and user discussion does create more context, which is useful for AI.

          However, I&#x27;d bring in Brooks discussion on essential and accidental complexity. In other words, there being No Silver Bullet <a href="https:&#x2F;&#x2F;www.cs.unc.edu&#x2F;techreports&#x2F;86-020.pdf" rel="nofollow">https:&#x2F;&#x2F;www.cs.unc.edu&#x2F;techreports&#x2F;86-020.pdf

          The problem with specification and user discussion is they still have errors that code has. But unlike code, there are no tests to confirm correctness.

          So now we have a definition of the system in a non-exact language with no ability to test to confirm it&#x27;s correctness. The code holds the essential complexity and now we are adding accidental complexity on top to manage.

          Again agree the specifications and user discussion provides context for the AI. However, a well written test suite provides similar context that can actually confirm correctness of the system.

          However, saying all the above. Focus of ImpactGate ( <a href="https:&#x2F;&#x2F;impactgate.officefloor.net" rel="nofollow">https:&#x2F;&#x2F;impactgate.officefloor.net ) is about erosion of the code, not correctness.

          1. visarga · · focus · HN ↗
            You usually don&#x27;t know what you want upfront, in real life it is a stream of specification and steering.
            1. sagenschneider · · focus · HN ↗
              Yes, agree. It&#x27;s a learning process. I tend to find when I build systems that at some point you need to stop analysis and just start building things to explore the problem. As you do, you prototype, refactor and possibly throw out ideas in favour of understanding the problem and discovery the real solution.

              The code becomes a reflection of that.

              I&#x27;m interested in your experiences of capturing specifications and user discussion on whether this captures the end intentions? Or whether it keeps you focused on earlier dead end directions?

              1. visarga · · focus · HN ↗
                I think agents are pretty capable of reading a log when given the explicit task of extracting the latest version of what the user wants. In general, they work well for direct tasks like this. They don&#x27;t forget and do something else the way they do when they are deep into development work or debugging.

                Besides intent, I also mine signs of &quot;user friction,&quot; which I use as input for the agent to come up with new tests. What I complain about is one of the signals driving testing.

                1. sagenschneider · · focus · HN ↗
                  Yes, agree very capable of consuming large amounts of information

                  I&#x27;d be interested to see what happens:

                  - to token counts after a year of so of changes, as the specification list grows?

                  - how it goes with concurrent changes in teams?

                  Plus whether asking AI to add good commenting to the code could achieve the same thing?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.