It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next (<a href="https://github.com/antirez/ds4/pull/1115" rel="nofollow">https://github.com/antirez/ds4/pull/1115 if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.
vlowther · · focus · HN ↗
ttoinou · · focus · HN ↗
vlowther · · focus · HN ↗
ttoinou · · focus · HN ↗