Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.
Depends on how many times you need to iterate to get the result you want. If you need to run the fast model 5 times to get the results you need, compared to 1-2 times for a smarter but slower model, you've just eroded any advantage that the speed gave you.
walrus01 · · focus · HN ↗
RussianCow · · focus · HN ↗
_aavaa_ · · focus · HN ↗
jgalt212 · · focus · HN ↗
RussianCow · · focus · HN ↗
So, as always: it depends on the use case.
jgalt212 · · focus · HN ↗