OpenAI's Misalignment Framework: A Tactical Bid to Preempt Global AI Governance
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
OpenAI's Misalignment Framework: A Tactical Bid to Preempt Global AI Governance
Unofficial Hacker News client; not affiliated with Y Combinator.
dcow · · focus · HN ↗
On one front it implies the model has a "mind of its own" (whether it does or not is besides the point). Why do we perceive human judgement as somehow more trustworthy than that of a model? I feel like I've experienced human misalignment somewhat regularly in life.
On another front I'm failing to conceptualize how alignment can be objective. How can you measure alignment when reasonable people will disagree whether actions are aligned or not? All the time I see humans operating in different zones of alignment with whatever goal they're trying to achieve and I suspect it's even a feature (socially) that we have people calibrated differently.
Do I want a model that's trying to push the boundaries of scientific understanding to be aligned strictly with the current dogmatic thinking? Or do I want it to "get creative" and think outside the box?
It seems to me more like accountability is the issue.
prerok · · focus · HN ↗
However, I have to say I also do not appreciate comparison that is continuously drawn with coworkers. As you say, it's a question of accountability but when the main agent will maliciously instruct the sub agents, whose fault is it then?
Yes, the person running this crap is at fault, not the CEO that's shoving it down their throat and definitely not the company that produced the AI.
Sorry for the rant, but seriously, if a person's goals do not align with the team's or company's we part ways. What do we do with AI? Stop using it?
dcow · · focus · HN ↗