I've been very excited with the most recent speed improvements for GLM5.3-Flash on DGX Spark clusters. It really feels close to what Opus ~4.5 was like to talk to. It's not quite there yet on consistency, but it's a really nice experience. Less guardrails and high quality abliterated versions further enhance its usefulness.
Though, Qwen3.8-Flash-Next is very close to that level while requiring fewer resources to run, so I'm really looking forward to Qwen4.
Yeah, I finally got a decent version of GLM 5.3 flash running on my dual sparks. It's not that fast (faster that sol 6.1 though, lol), but it's manageable. Qwen 3.8 on the other hand, that thing really performs on dual sparks. You can get consistent multiple (3-4) concurrent sessions at 40+ tokens per second. It's pretty nice. Good progress, but definitely not that cost effective. I mean, OPUS 5.5 has been pretty magnificent.
But nothing compares to when 4.5 came out. I still remember, it was just about Thanksgiving and that just was awesome.
hgoel · · focus · HN ↗
Though, Qwen3.8-Flash-Next is very close to that level while requiring fewer resources to run, so I'm really looking forward to Qwen4.
vardalab · · focus · HN ↗