Pi is so good, for both running Qwen 3.8 Flash Next locally at home, and for using all the models available at my work. Shockingly useful, fast, it's TUI doesn't suck (unlike my work's own agent CLI: it's really powerful, but man that actual TUI itself is a bit rough, its too GUI-like), and extensible.
Adding MCP support is lovely, codemode sounds super interesting, and I'm super excited to take advantage of it. Now it's 1.0 I'm hoping I can convince IT to let us use it officially.
A DGX Spark-like, the Asus GX10. I’m kicking myself that I didn’t buy a second one when I thought about it months ago, but Nvidia’s NVFP4 quantisation of Flash Next and offloading the ngram table to NVMe has worked well
I run Qwen3.8 Flash Next on 2x Sparks at $WORK and it has been great! It can run 16 concurrent sessions with decent speed (~30-40 t/s per session) and can hold 8 full 256k contexts in cache. It has turned out to be a great production setup for our company.
I have a half finished repo that sets it all up for production (using systemd, not hacky scripts and one-off docker commands), if someone is interested in this, tell me, and it might give me the push to actually finish up the latest threads and publish it!
girvo · · focus · HN ↗
Adding MCP support is lovely, codemode sounds super interesting, and I'm super excited to take advantage of it. Now it's 1.0 I'm hoping I can convince IT to let us use it officially.
Zambyte · · focus · HN ↗
girvo · · focus · HN ↗
rsolva · · focus · HN ↗
I have a half finished repo that sets it all up for production (using systemd, not hacky scripts and one-off docker commands), if someone is interested in this, tell me, and it might give me the push to actually finish up the latest threads and publish it!