By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
I also no longer trust benchmarks on this one.
When the context gets a bit longer and the problem harder low quant models often produce worse output for me.
Sometimes they even loop.
Interestingly different formats also often behave differently.
GGUF unsloth is so far the best for me.
infogulch · · focus · HN ↗
If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
kadushka · · focus · HN ↗
mixermachine · · focus · HN ↗
Interestingly different formats also often behave differently. GGUF unsloth is so far the best for me.