Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.
Is it possible to construct a control system where bad, fast and cheap can become good, fast, and cheap through repeated sampling and a strong spec/eval harness?
I am trying to keep an open mind with AI, but I also have little understanding of control theory, trying to learn.
You can, but you need to break the problem into much smaller tasks, then check those answers, and finally have a harness that handles all the context, task breakup, task definitions, and validations each round.
Depends on how many times you need to iterate to get the result you want. If you need to run the fast model 5 times to get the results you need, compared to 1-2 times for a smarter but slower model, you've just eroded any advantage that the speed gave you.
walrus01 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
RussianCow · · focus · HN ↗
_aavaa_ · · focus · HN ↗
hermannj314 · · focus · HN ↗
I am trying to keep an open mind with AI, but I also have little understanding of control theory, trying to learn.
freakynit · · focus · HN ↗
calgoo · · focus · HN ↗
RussianCow · · focus · HN ↗
jgalt212 · · focus · HN ↗
RussianCow · · focus · HN ↗
So, as always: it depends on the use case.
jgalt212 · · focus · HN ↗