Once Claude can measure something, it can make it faster
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Once Claude can measure something, it can make it faster
Unofficial Hacker News client; not affiliated with Y Combinator.
hungryhobbit · · focus · HN ↗
I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?
When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it
That's the entirety of Anthropic's billions of dollars of research: any prompt with the word "reasoning" is trying to hack Claude to figure out how it reasons!
A model like that should never have gotten out of QA, let alone been released.
railgunmerlin · · focus · HN ↗
Marciplan · · focus · HN ↗
[dead]
cyanydeez · · focus · HN ↗
might as well offer your life to a king to work in their fields.
hungryhobbit · · focus · HN ↗
Opus 5.5 literally refused to work OR EVEN TELL ME WHAT I'D "SAID" when it read that prompt.
Nothing to do with skill or the user at all: same exact prompt, three different models ... two worked, one didn't.
bitpush · · focus · HN ↗
This blogpost is about frontend performance. It'll be akin to you commenting on a swift blogpost saying 'How about Airpods noise cancellation'. Sure both are Apple, but they are wildly different teams.
hungryhobbit · · focus · HN ↗
In such a forum, it seems to me like it's fair game to point out that the company patting itself on the back about how great they are at programming (as evidenced in the article above about their 3x speed improvement) ...
... can't even make their latest model handle basic English without refusing to work.
jfidjcjwjcjwjd · · focus · HN ↗
s3p · · focus · HN ↗
In a blog post about Claude, i find it strange you get upset when people talk about Claude
[deleted] · · focus · HN ↗
[deleted]
s3p · · focus · HN ↗
In both instances, it's a side point that is actually tangentially related to the first. Not completely unrelated as you are implying
frumplestlatz · · focus · HN ↗
Every single time it triggered, it was due to a prompt written by their own model in a dynamic workflow. The self-serving nanny oversight has to go.
The fact that they label model distillation as an “attack” is genuinely hilarious after they “distilled“ their models from all of our work, and continue to do so.
I believe AI is here to stay and an incredibly powerful tool, but these companies, and especially Dario and Altman, are the very last people I want to see in charge of it.
copperx · · focus · HN ↗
They distilled all digitized human knowledge and artifacts and they're now complaining about someone copying their outputs saying it's a "national security concern."
I'm not sure about how to classify that. Hilarious? Pathetic? Sad? Hypocritical? Hyperdramatic? All of the above?
vikramkr · · focus · HN ↗
post-it · · focus · HN ↗
Did it explain it did it hallucinate?
hungryhobbit · · focus · HN ↗
It's more or less the same mistake we've seen Anthropic make repeatedly with it's brain-dead regex-based Fable/Mythos gates.
post-it · · focus · HN ↗
How do you know? How would the stupider model know?
nfcampos · · focus · HN ↗
dolmen · · focus · HN ↗
ronsor · · focus · HN ↗
PunchyHamster · · focus · HN ↗
And the goal is to slow competition down enough before IPO
crooked-v · · focus · HN ↗
rendaw · · focus · HN ↗