By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
Quantization is a category error, the thing you care about is not in weight space, so there’s unbounded error introduced by doing it. The thing you actually want to preserve is the knowledge manifold, but that is in a different vector space. Until we have some better understanding of how to interact with that space directly, rather than inferring it through distillation of reasoning traces, I would not anticipate truly low bit models to be useful.
infogulch · · focus · HN ↗
If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
kadushka · · focus · HN ↗
FuckButtons · · focus · HN ↗