‹ BackHN Continuity

Thread

What would a serious AI product look like?

176 points · 84 comments · lumpa

  1. ake2l · · focus · HN ↗
    For me the interesting part is not whether the model gets better. It probably will. The problem is when the same system creates the change, explains why the change is correct, and basically also grades itself. “Just review it carefully” does not scale either. After 50 correct looking changes humans start trusting the green output. I do too. I am experimenting with moving more of this outside the agent ... deterministic checks, frozen behaviour, architecture constraints, explicit evidence. And probably most important -> UNKNOWN when I simply cannot prove something. I increasingly think this is the missing layer in serious agentic engineering. Not another smarter agent judging the first agent, but boring independent machinery which does not care how convincing the explanation sounds.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.