I've been extensively using GLM 5.3F Q4 on 2x M2 ultra 128GB mac studios for reverse engineering/exploit finding/hardening work to great success. 50t/s tg and 600 t/s pp is more than enough for me.
Hopefully the Anthropic fearmongering doesn't stop/delay the 5.4 release.
it's a private fork of ds4 - vibecoded to optimize for my platform while maintaining correctness. There's a ton of headroom on the software side for consumer hardware.
especially with 5.3 Flash's combination of KDA+DSA attention, the decode speed scales amazingly well with context.
Very cool, congrats on getting it to work. Just like the good old days: hardware limitations stimulate creativity, and in many ways this is the new frontier: the democratization of this tech. Anthropic and OpenAI would love to be the new IBM/Microsoft/Google but I think their cycle of ascent and descent will be a lot shorter than those other three (and those cycles were getting shorter anyway).
Agreed - I think we're one or two (GLM/Deepseek probably)release cycles away from exactly the inflection point in the cycle you mention.
Once a certain baseline capable model is open and available (hardware non-withstanding, I know a 128 mac/spark is expensive now, but they don't need to get faster - just cheaper), there's no putting the toothpaste back in the tube (I hope).
It is incredible how fast these open models are improving, I think Anthropic really messed up here, they missed their window of opportunity for an IPO because they got greedy. In the last three months there have been a whole raft of major open model releases and it does not look as if they're slowing down, besides that, the smaller models are getting more and more capable.
sdlkj- · · focus · HN ↗
Hopefully the Anthropic fearmongering doesn't stop/delay the 5.4 release.
jacquesm · · focus · HN ↗
sdlkj- · · focus · HN ↗
especially with 5.3 Flash's combination of KDA+DSA attention, the decode speed scales amazingly well with context.
jacquesm · · focus · HN ↗
sdlkj- · · focus · HN ↗
Once a certain baseline capable model is open and available (hardware non-withstanding, I know a 128 mac/spark is expensive now, but they don't need to get faster - just cheaper), there's no putting the toothpaste back in the tube (I hope).
jacquesm · · focus · HN ↗