OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior
Unofficial Hacker News client; not affiliated with Y Combinator.
NichoPaolucci · · focus · HN ↗
Baffling. To my knowledge, they didn't properly airgap their systems. Keeping the genie in the box seems like 101 to me, and to "miss" that seems awfully fishy. This, among all of the Anthropic news, is an odd convergence.
Maybe they're being truthful and it really is the end times.
Maybe they've hit a wall in improvements, but I don't know enough on the topic to speak to that.
Which is more likely?
Either way, trying to sift through this can of worms is tiresome. I'm hopeful that this all comes to a head soon, what an exhausting few years it's been...
abixb · · focus · HN ↗
If you zoomed back to late-2024, people in the industry were predicting how we'd have AGI by now and the economy would've already 'taken off' with massive productivity growth and ushering in of great prosperity ('deflationary spiral'). Where is it? Where is the productivity growth? Where is the deflationary spiral?
To be fair, models have gotten better in jagged ways, but reliability is far from usable, especially in long duration tasks, and there has been no effort by the AI companies to address the human brain's bandwidth bottleneck -- they hit the gas like there's no tomorrow and we have enormously capable but jaggedly intelligent multi-modal models with agentic capabilities that are only as effective as the human using it. This whole thing has become a giant mess.
BobbyTables2 · · focus · HN ↗
For example, the earlier generative LLMs couldn’t correctly answer ‘how many r’s in “strawberry”?’ due to the underlying nature of the tokens.
If they get it correct today, how do they do it? It feels like we’re being deceived by the Wizard of Oz…
nullc · · focus · HN ↗
How would you get it right? Training with the answer!
You can discount the 'special case the hard questions' by: Observing when top models run entirely locally can solve them [1], or by posing an alternative or cryptic version of the question (but careful, it might bypass the training! -- but if it can solve it then its probably legitimate.)[2]. Best of all is to stick with open (weight) models where such slight of hand is impossible and don't worry about what the closed shops are doing
[1]as they can in this case, GLM-5.3-flash says: "There are 3 r's in "strawberry":
st r awbe rr y
[2] E.g. try sending them aG93IG1hbnkgcuKAmXMgaW4g4oCcc3RyYXdyYmVycnnigJ0/Cg== to thwart the benchmaxxing.