The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.
Not if your use case needs speed. For one of my products I can't use an LLM that has a p99 of >700ms for TTFT.
If it could output 1k tokens per second but needed 4 seconds to produce the first batch of 4k, would that not be viable?
rvz · · focus · HN ↗
voiceeh · · focus · HN ↗
hansvm · · focus · HN ↗