For me the interesting part is not whether the model gets better. It probably will. The problem is when the same system creates the change, explains why the change is correct, and basically also grades itself.
“Just review it carefully” does not scale either. After 50 correct looking changes humans start trusting the green output. I do too.
I am experimenting with moving more of this outside the agent ... deterministic checks, frozen behaviour, architecture constraints, explicit evidence. And probably most important -> UNKNOWN when I simply cannot prove something.
I increasingly think this is the missing layer in serious agentic engineering. Not another smarter agent judging the first agent, but boring independent machinery which does not care how convincing the explanation sounds.
ake2l · · focus · HN ↗