I'm very thankful for the section on reproducibility. I argue this is the single biggest hangup for the entire space. You CAN have temperature and determinism. I've been waiting for six years for a major provider to offer it, there is demand, but I've slowly come to realize the current game theory does not support it.
For providers, not supporting deterministic eval means:
- users use more tokens = more money
- providers can generate more tokens per compute = more money
- providers have cheaper hardware options (GPUs) = more money
- providers models are harder to extract/distill = more money
- providers are harder to hold liable for outputs = more money
- providers can secretly use other models = more money
- providers are harder to compare against others = more money
- providers can cherry pick performance results = more money
This is true, but there's also the cases where slight differences in prompt yield wildly different results. In any programming language, if I add a clause to a conditional like "if car is red or car is blue", that behaves predictably--and if it doesn't we can dig into the debugger, assembly, etc. If I do that with an LLM, that can change everything, and there's no way to "debug" it.
This kind of thing (plus the cost) really limits what they can realistically be used for. A lot of things are tolerant of even lots of fuzziness (suggestions you can ignore, work you can redo, etc), but that subset of applications doesn't justify the boggling capital investment or the ongoing compute needs.
So, my guess is we're probably in for a couple more years of discovering what these models are good for. Coding: meh, kinda. Hacking: wow amazing. Writing a novel: no. Reviewing your work: incredible. And so it goes. This is probably what pops the bubble: we find the small subset of applications this stuff is useful for, and then it's a bag holding race.
ramity · · focus · HN ↗
For providers, not supporting deterministic eval means:
- users use more tokens = more money
- providers can generate more tokens per compute = more money
- providers have cheaper hardware options (GPUs) = more money
- providers models are harder to extract/distill = more money
- providers are harder to hold liable for outputs = more money
- providers can secretly use other models = more money
- providers are harder to compare against others = more money
- providers can cherry pick performance results = more money
camgunz · · focus · HN ↗
This kind of thing (plus the cost) really limits what they can realistically be used for. A lot of things are tolerant of even lots of fuzziness (suggestions you can ignore, work you can redo, etc), but that subset of applications doesn't justify the boggling capital investment or the ongoing compute needs.
So, my guess is we're probably in for a couple more years of discovering what these models are good for. Coding: meh, kinda. Hacking: wow amazing. Writing a novel: no. Reviewing your work: incredible. And so it goes. This is probably what pops the bubble: we find the small subset of applications this stuff is useful for, and then it's a bag holding race.