I don't want to bash on the article, but a much more interesting take with llm that worked at 1kk TPS would the the countless amount of real time applications you could build with it.
I like electronic music, I dance to it a lot. What if I had a LLM that had enough throughput and low latency to do the reverse, take my dancing and generate music.
We are still stuck stargazing raw artificial cognitive intelligence, but real progress will be when we stop perceiving these as external intelligence and just an extension of ourselves.
I think you are right and I think we are close to that speed, with codex ultrafast we are near 500 TPS but only with Astra and that is not financially viable. I would guess we will have wide spread near 1000 TPS with 6.1 sol level model with enough usage to, without worry, do the kind of things you describe in about 6 months max
gchamonlive · · focus · HN ↗
I like electronic music, I dance to it a lot. What if I had a LLM that had enough throughput and low latency to do the reverse, take my dancing and generate music.
We are still stuck stargazing raw artificial cognitive intelligence, but real progress will be when we stop perceiving these as external intelligence and just an extension of ourselves.
echohive42 · · focus · HN ↗