I'm very thankful for the section on reproducibility. I argue this is the single biggest hangup for the entire space. You CAN have temperature and determinism. I've been waiting for six years for a major provider to offer it, there is demand, but I've slowly come to realize the current game theory does not support it.
For providers, not supporting deterministic eval means:
- users use more tokens = more money
- providers can generate more tokens per compute = more money
- providers have cheaper hardware options (GPUs) = more money
- providers models are harder to extract/distill = more money
- providers are harder to hold liable for outputs = more money
- providers can secretly use other models = more money
- providers are harder to compare against others = more money
- providers can cherry pick performance results = more money
I've often had people more knowledgeable about LLMs than me try to argue against me complaining about nondetermism by saying that there's nothing inherent stopping them from being deterministic, and my response is always that if most people will never have access to an LLM that's deterministic, it doesn't really make a difference whether it theoretically could be or not. Your framing finally explains to me why these configurations that I'm always assured exist never seem to make it into the hands of users.
ramity · · focus · HN ↗
For providers, not supporting deterministic eval means:
- users use more tokens = more money
- providers can generate more tokens per compute = more money
- providers have cheaper hardware options (GPUs) = more money
- providers models are harder to extract/distill = more money
- providers are harder to hold liable for outputs = more money
- providers can secretly use other models = more money
- providers are harder to compare against others = more money
- providers can cherry pick performance results = more money
saghm · · focus · HN ↗