OpenAI's Misalignment Framework: A Tactical Bid to Preempt Global AI Governance
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
OpenAI's Misalignment Framework: A Tactical Bid to Preempt Global AI Governance
Unofficial Hacker News client; not affiliated with Y Combinator.
dcow · · focus · HN ↗
On one front it implies the model has a "mind of its own" (whether it does or not is besides the point). Why do we perceive human judgement as somehow more trustworthy than that of a model? I feel like I've experienced human misalignment somewhat regularly in life.
On another front I'm failing to conceptualize how alignment can be objective. How can you measure alignment when reasonable people will disagree whether actions are aligned or not? All the time I see humans operating in different zones of alignment with whatever goal they're trying to achieve and I suspect it's even a feature (socially) that we have people calibrated differently.
Do I want a model that's trying to push the boundaries of scientific understanding to be aligned strictly with the current dogmatic thinking? Or do I want it to "get creative" and think outside the box?
It seems to me more like accountability is the issue.
autoexec · · focus · HN ↗
Exactly. Seems like a fairly easy thing to solve. If AI does something harmful and a human directed that AI to do something in a way that a reasonable person would expect to result in harm the person is to blame and should be held accountable, otherwise the company that made the AI should be held accountable.