This doesn't seem like a well-thought-out post. I mean, I probably could have told you that from the fact that it's either pro-AI or anti-AI (in this case the latter), but this is a particularly poorly thought out post. Others here have pointed out the water-wasting red herring, but there are several other pieces of nonsense here — the general environmental doom-mongering, and labeling recursive self-improvement as "mystical", for example.
Me, I pointed Claude Opus 5 at a C codebase I've been working on for years and it immediately found five serious bugs and told me how to fix them, as well as how to use Linux system calls I wasn't familiar with to solve some other problems. And I gave it a new language for PEGs I'd written up a couple of years ago, and it wrote me a working implementation in OCaml that afternoon, fixing several bugs in my example grammars along the way and resolving a conceptual problem I'd had a roadblock on. I asked Fable 5.1 to implement a minimal proof assistant, and it wrote a clone of J-Bob in Python, which I'm studying now. Although there is a certain firehose quality to this stuff, I sure don't intend to use AI to actively de-skill my brain.
That doesn't guarantee that AI will be a beneficial innovation overall, of course! As with most innovations, it probably depends on the balance between people being able to use the innovation to gain control over their own lives, and people being able to use it to gain control over others' lives. (Barring a FOOM scenario, of course, where our prediction ability is nonexistent.)
So, I don't think a pro-AI post can be well thought out, either. It's too early and chaotic to understand what's going to happen, and it may actually be uncertain. Pro and anti are both far too simplistic.
My impression was that the frontier AI coding models are kind of resetting the barriers to entry/reward ratios to before the dot com boom, when people didn't go into this career for money, and so most folks that ended up in it were the 1% kind of talents who really wanted it. Perhaps that's an elitist thing to say? I don't know. That's not to say that AI is pushing people away from computing, it's just that it's creating such massive low resistance paths to getting to the goal without the "productive friction."
p.s. and at the same time, it's such a fantastic learning tool. It distills the intuition of a whole world and is able to transmit it on demand, like in your examples. It's an interesting dichotomy.
> It distills the intuition of a whole world and is able to transmit it on demand, like in your examples. It's an interesting dichotomy.
The only problem I have at the moment is it's very prone to hallucinations, even now, yes frontier models astra fable with all the bells and whistles. This makes it hard to confidently use while learning because I have to be on the lookout for lies while I'm learning, which is precisely the moment I am least able to distinguish them. So instead theres just a constant low lying dread.
Nonetheless, I am able to get some value out of them. Just not all, everything requires manual effort to duplicate and check which you should probably be doing anyway as part of learning
I mentioned the minimal proof assistant I elicited last night. When Fable 5.1 thought it was done, I asked how we knew the prover was sound (which means, in the jargon, that the theorems that it proves are actually true). It checked and immediately found five different ways it could "prove" false theorems with it (and fixed them).
Then it suggested writing a fuzzer to try to flush out more soundness problems. It loves fuzzers! And usually they are an excellent cost/benefit tradeoff. However, in this case, I questioned whether a fuzzer would ever actually succeed at finding proofs of a theorem it was set to prove, even unsound proofs, and after doing some tests it admitted that the fuzzer it had proposed would have been completely useless.
Then the conversation was incorrectly flagged as me working on a malicious security attack, so I was downgraded to Opus 4.8. I switched back to Fable, renamed the file in its scratchpad, and asked it to please use the term "generative testing" instead of "fuzzing". Thanks to not using the computer-security name for generative testing, there were no further false flags.
Then, this morning, I was reading the spec it proved its program fulfilled, and asked whether a certain trivially incorrect alternative program would also fulfill the spec (a misformalization problem rather than a soundness problem). It churned away for a while, discovering that while, actually, no, that program would be rejected, a different trivially incorrect program would pass, and credited me in the docs with pointing out the issue. I pointed out that in fact the issue it had found was completely different from my stupid misreading of the spec. It fixed the doc.
Fable 5.1 isn't Mythos but it's generally considered to be a "frontier model".
So, from my point of view, the whole experience has been a constant fractal of hallucination, in which I have to constantly struggle to keep my grip (and Claude's grip) on actual reality, because it's so willing to make up surface-plausible nonsense.
______
P.S. Also, in another task today, it thought pip wasn't installed and was trying to figure out how to work around it. But that's not so much a hallucination as a failure to recheck assumptions — I hadn't installed pip on the machine before the first Claude work on it, and it just assumed that was still the case. Also I think that might have been Opus rather than Fable, so it's not as strong a case.
I feel like this pattern is a great way to practice and sharpen critical thinking. It's like a whole new skill to deal with computing systems in this way.
Not the GP, and this is just the free ChatGPT, not a frontier model, but just a few days ago it happily confabulated an entirely incorrect version of the plot of Iain M. Banks’s "Matter" when I asked it to analyze the novel in a certain context. Claude fared better, though it also made some mistakes. I don’t expect the models to know or recall the specifics of the plot of every novel out there, but it would be nice if they didn’t make things up.
Probably over twenty today in my work day alone. Its a normal part of working with agents. They do not always have the context they need, but fail to be aware of this.
Misinterpreting the results of performance testing it had just run to argue for the exact opposite of what the data it generated showed.
Stating something was current guidance (a quick read shows the document it found was from 2016, it was served the last edit date in it's API call, along with newer documents that contradict it).
Assuming what a Jira ticket said without ever reading it with it's tools and then doubling down on the contents it has assumed.
On learning specifically, it constantly gets grammar in foreign languages wrong it can write fluently when prompted correctly but when asked the sort of wrongheaded questions from flawed premises learners commonly make it is prone to making stuff up. I experienced this yesterday when asking about how danish comparisons work.
Heck i mean see <a href="https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/" rel="nofollow">https://alignment.openai.com/misalignment-reports/self-gener...
They are prone to make up restrictions for themselves you never asked for. Ive seen this behavior on occasion also.
lol i love playing with food labels. the things it spits out for anything de papa in an otherwise English paragraph is terrifying. ai must solve the spanglish potato problem. also the spanglish con issue.
kragen · · focus · HN ↗
Me, I pointed Claude Opus 5 at a C codebase I've been working on for years and it immediately found five serious bugs and told me how to fix them, as well as how to use Linux system calls I wasn't familiar with to solve some other problems. And I gave it a new language for PEGs I'd written up a couple of years ago, and it wrote me a working implementation in OCaml that afternoon, fixing several bugs in my example grammars along the way and resolving a conceptual problem I'd had a roadblock on. I asked Fable 5.1 to implement a minimal proof assistant, and it wrote a clone of J-Bob in Python, which I'm studying now. Although there is a certain firehose quality to this stuff, I sure don't intend to use AI to actively de-skill my brain.
That doesn't guarantee that AI will be a beneficial innovation overall, of course! As with most innovations, it probably depends on the balance between people being able to use the innovation to gain control over their own lives, and people being able to use it to gain control over others' lives. (Barring a FOOM scenario, of course, where our prediction ability is nonexistent.)
So, I don't think a pro-AI post can be well thought out, either. It's too early and chaotic to understand what's going to happen, and it may actually be uncertain. Pro and anti are both far too simplistic.
foobarian · · focus · HN ↗
p.s. and at the same time, it's such a fantastic learning tool. It distills the intuition of a whole world and is able to transmit it on demand, like in your examples. It's an interesting dichotomy.
RugnirViking · · focus · HN ↗
The only problem I have at the moment is it's very prone to hallucinations, even now, yes frontier models astra fable with all the bells and whistles. This makes it hard to confidently use while learning because I have to be on the lookout for lies while I'm learning, which is precisely the moment I am least able to distinguish them. So instead theres just a constant low lying dread.
Nonetheless, I am able to get some value out of them. Just not all, everything requires manual effort to duplicate and check which you should probably be doing anyway as part of learning
fragmede · · focus · HN ↗
kragen · · focus · HN ↗
Then it suggested writing a fuzzer to try to flush out more soundness problems. It loves fuzzers! And usually they are an excellent cost/benefit tradeoff. However, in this case, I questioned whether a fuzzer would ever actually succeed at finding proofs of a theorem it was set to prove, even unsound proofs, and after doing some tests it admitted that the fuzzer it had proposed would have been completely useless.
Then the conversation was incorrectly flagged as me working on a malicious security attack, so I was downgraded to Opus 4.8. I switched back to Fable, renamed the file in its scratchpad, and asked it to please use the term "generative testing" instead of "fuzzing". Thanks to not using the computer-security name for generative testing, there were no further false flags.
Then, this morning, I was reading the spec it proved its program fulfilled, and asked whether a certain trivially incorrect alternative program would also fulfill the spec (a misformalization problem rather than a soundness problem). It churned away for a while, discovering that while, actually, no, that program would be rejected, a different trivially incorrect program would pass, and credited me in the docs with pointing out the issue. I pointed out that in fact the issue it had found was completely different from my stupid misreading of the spec. It fixed the doc.
Fable 5.1 isn't Mythos but it's generally considered to be a "frontier model".
So, from my point of view, the whole experience has been a constant fractal of hallucination, in which I have to constantly struggle to keep my grip (and Claude's grip) on actual reality, because it's so willing to make up surface-plausible nonsense.
______
P.S. Also, in another task today, it thought pip wasn't installed and was trying to figure out how to work around it. But that's not so much a hallucination as a failure to recheck assumptions — I hadn't installed pip on the machine before the first Claude work on it, and it just assumed that was still the case. Also I think that might have been Opus rather than Fable, so it's not as strong a case.
foobarian · · focus · HN ↗
Sharlin · · focus · HN ↗
inquirerGeneral · · focus · HN ↗
[dead]
RugnirViking · · focus · HN ↗
Misinterpreting the results of performance testing it had just run to argue for the exact opposite of what the data it generated showed.
Stating something was current guidance (a quick read shows the document it found was from 2016, it was served the last edit date in it's API call, along with newer documents that contradict it).
Assuming what a Jira ticket said without ever reading it with it's tools and then doubling down on the contents it has assumed.
On learning specifically, it constantly gets grammar in foreign languages wrong it can write fluently when prompted correctly but when asked the sort of wrongheaded questions from flawed premises learners commonly make it is prone to making stuff up. I experienced this yesterday when asking about how danish comparisons work.
Heck i mean see <a href="https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/" rel="nofollow">https://alignment.openai.com/misalignment-reports/self-gener...
They are prone to make up restrictions for themselves you never asked for. Ive seen this behavior on occasion also.
jambalaya8 · · focus · HN ↗
therealdrag0 · · focus · HN ↗