‹ BackHN Continuity

Thread

What would a serious AI product look like?

176 points · 84 comments · lumpa

  1. ramity · · focus · HN ↗
    I'm very thankful for the section on reproducibility. I argue this is the single biggest hangup for the entire space. You CAN have temperature and determinism. I've been waiting for six years for a major provider to offer it, there is demand, but I've slowly come to realize the current game theory does not support it.

    For providers, not supporting deterministic eval means:

    - users use more tokens = more money

    - providers can generate more tokens per compute = more money

    - providers have cheaper hardware options (GPUs) = more money

    - providers models are harder to extract/distill = more money

    - providers are harder to hold liable for outputs = more money

    - providers can secretly use other models = more money

    - providers are harder to compare against others = more money

    - providers can cherry pick performance results = more money

    1. augment_me · · focus · HN ↗
      Very good points. Incentives are just terrible for this.

      Add in:

      - harder to audit

      - move cost of failure/reprompts to the user

      - kind of noted by you, but all kinds of quantization, model pruning, model routing, A/B testing becomes invisible and without any repercussions. The ways to cost-optimize are just crazy.

      IIRC Thinking Machines had a mode with deterministic numerics but it's more expensive to run due to limitations this imposes on cross-batch ops and ordering of floating point reductions, and their model is not great overall.

    2. camgunz · · focus · HN ↗
      This is true, but there's also the cases where slight differences in prompt yield wildly different results. In any programming language, if I add a clause to a conditional like "if car is red or car is blue", that behaves predictably--and if it doesn't we can dig into the debugger, assembly, etc. If I do that with an LLM, that can change everything, and there's no way to "debug" it.

      This kind of thing (plus the cost) really limits what they can realistically be used for. A lot of things are tolerant of even lots of fuzziness (suggestions you can ignore, work you can redo, etc), but that subset of applications doesn't justify the boggling capital investment or the ongoing compute needs.

      So, my guess is we're probably in for a couple more years of discovering what these models are good for. Coding: meh, kinda. Hacking: wow amazing. Writing a novel: no. Reviewing your work: incredible. And so it goes. This is probably what pops the bubble: we find the small subset of applications this stuff is useful for, and then it's a bag holding race.

    3. saghm · · focus · HN ↗
      I've often had people more knowledgeable about LLMs than me try to argue against me complaining about nondetermism by saying that there's nothing inherent stopping them from being deterministic, and my response is always that if most people will never have access to an LLM that's deterministic, it doesn't really make a difference whether it theoretically could be or not. Your framing finally explains to me why these configurations that I'm always assured exist never seem to make it into the hands of users.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.