This story is written like it vindicates AI medicine but really it seems like the human medical system dropped the ball on the MRI. ChatGPT didn't have any brilliant insight, it just (eventually) suggested the same thing a human doctor had already suggested, but this time the author and his wife pushed harder for it to happen.
To use the metaphor from the article, they found the keys in the last place they looked.
this is one case. there's no proof here that AI will generally do better than human doctors. consider also that the power it has to convince people to push for better medical treatments can also be used to push for worse ones, or to halt standard treatment entirely, at which point it is much harder for the medical system to intervene as they aren't interacting with doctors.
There is some evidence that has been collected so far that, taken across large sample sizes, good AI systems are now better at diagnosing certain diseases, etc…. We started here with Mycin and e-Mycin back in the 60s. Hopefully, these systems can now start to be better integrated into modern medicine systems.
However, the real problem is that modern AI systems still hallucinate. And there is not yet any evidence that I’ve seen that AI systems can be trusted to work at these levels of efficiency and correctness at the tactical single-patient level, as opposed to the strategic population level.
It feels like moving toward a local maxima. In this example the LLM advice outperformed one doctor who dropped the ball in some ways, but I do not look forward to a world where we make it even more difficult to talk to doctors, and route most medical interactions to LLMs. I expect that in 5 years I'll be futilely shouting "representative" into my phone as my appendix ruptures.
With this sort of advice from LLMs, I think there's a lot of selective recall that's easily overlooked -- where we gloss over all the bad, dead-end lines of advice we seek from the LLM oracle.
ForHackernews · · focus · HN ↗
To use the metaphor from the article, they found the keys in the last place they looked.
saulpw · · focus · HN ↗
peesem · · focus · HN ↗
bradknowles · · focus · HN ↗
However, the real problem is that modern AI systems still hallucinate. And there is not yet any evidence that I’ve seen that AI systems can be trusted to work at these levels of efficiency and correctness at the tactical single-patient level, as opposed to the strategic population level.
I think lots more work needs to be done here.
nerevarthelame · · focus · HN ↗
With this sort of advice from LLMs, I think there's a lot of selective recall that's easily overlooked -- where we gloss over all the bad, dead-end lines of advice we seek from the LLM oracle.