Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.
Reminds me of how slot machine users swear the odds have changed on a machine.
also when someone says you just have to prompt it a certain way it reminds me of people who think they can get better results out of a slot machine by pressing buttons in a certain order
The providers of these models also design the UX similarly to slot machines (run it x amount of times for better results, multiplying your spend) this isnt a coincidence and they're playing into the gambler mentality, and probably hire UX designers that specialize in this.
Wtf are you on my dude. Anthropic UI is designed like a slot machine? Hiring slot machine specialists? Sometimes I can’t believe im even on HN anymore with comments like this.
I think some of it comes from that they do not publicly let you see the random seed. So each time you ask the answer is different (like a slot machine) and if they let users use the random seed it would let people much more accurately assess if an underlying model changed somehow (same seed and same input will always have the same output).
Of course the closed Anthropic would never share this, it would definitely take away the 'magic' feeling of the AI
To my understanding, with batched inference and other "optimizations" you wouldn't get the exact same token predictions even with temp=0.0.
With models there are a bunch of other dials that can be tuned even if the model itself remains exactly the same.
Are those dials set the same across all hardware configurations and clusters? Does model behavior average out the same across different hardware?
There are just too many different buttons that can be set to really trust a provider either not to directly commit fraud, or indirectly commit fraud with system complexity affecting the output.
mlmonkey · · focus · HN ↗
physicallyIllfr · · focus · HN ↗
also when someone says you just have to prompt it a certain way it reminds me of people who think they can get better results out of a slot machine by pressing buttons in a certain order
The providers of these models also design the UX similarly to slot machines (run it x amount of times for better results, multiplying your spend) this isnt a coincidence and they're playing into the gambler mentality, and probably hire UX designers that specialize in this.
cheevly · · focus · HN ↗
ddxv · · focus · HN ↗
Of course the closed Anthropic would never share this, it would definitely take away the 'magic' feeling of the AI
mh- · · focus · HN ↗
pixl97 · · focus · HN ↗
Are those dials set the same across all hardware configurations and clusters? Does model behavior average out the same across different hardware?
There are just too many different buttons that can be set to really trust a provider either not to directly commit fraud, or indirectly commit fraud with system complexity affecting the output.