ImpactGate: A merge gate that scores the structural decay AI adds
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
ImpactGate: A merge gate that scores the structural decay AI adds
Unofficial Hacker News client; not affiliated with Y Combinator.
MichaelNolan · · focus · HN ↗
sagenschneider · · focus · HN ↗
visarga · · focus · HN ↗
Compactly formatted user messages are something an agent can ingest in a few minutes, even if they are thousands of lines long. And the quality of those messages is great: they don't track what the agent does well, only what changes and what breaks.
Having this top-down view helps a lot. Usually, within a session and deep into a task, the agent loses the global perspective and optimizes for local success. I find it weird there is no harness that treats user messages as high value signal (except my own, of course, I have it, <a href="https://github.com/horiacristescu/playbook-harness" rel="nofollow">https://github.com/horiacristescu/playbook-harness).
sagenschneider · · focus · HN ↗
However, I'd bring in Brooks discussion on essential and accidental complexity. In other words, there being No Silver Bullet <a href="https://www.cs.unc.edu/techreports/86-020.pdf" rel="nofollow">https://www.cs.unc.edu/techreports/86-020.pdf
The problem with specification and user discussion is they still have errors that code has. But unlike code, there are no tests to confirm correctness.
So now we have a definition of the system in a non-exact language with no ability to test to confirm it's correctness. The code holds the essential complexity and now we are adding accidental complexity on top to manage.
Again agree the specifications and user discussion provides context for the AI. However, a well written test suite provides similar context that can actually confirm correctness of the system.
However, saying all the above. Focus of ImpactGate ( <a href="https://impactgate.officefloor.net" rel="nofollow">https://impactgate.officefloor.net ) is about erosion of the code, not correctness.
visarga · · focus · HN ↗
sagenschneider · · focus · HN ↗
The code becomes a reflection of that.
I'm interested in your experiences of capturing specifications and user discussion on whether this captures the end intentions? Or whether it keeps you focused on earlier dead end directions?
visarga · · focus · HN ↗
Besides intent, I also mine signs of "user friction," which I use as input for the agent to come up with new tests. What I complain about is one of the signals driving testing.
sagenschneider · · focus · HN ↗
I'd be interested to see what happens:
- to token counts after a year of so of changes, as the specification list grows?
- how it goes with concurrent changes in teams?
Plus whether asking AI to add good commenting to the code could achieve the same thing?