Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
If you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
om8 · · focus · HN ↗
om8 · · focus · HN ↗