By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
Not an expert, but doesn’t that produce lower quality results, the same way a 1MP image isn’t lower quality than a 20mp image downscaled to 1MP? (Everything else equal)
infogulch · · focus · HN ↗
If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
kadushka · · focus · HN ↗
montroser · · focus · HN ↗
brookst · · focus · HN ↗