Ternary LLMs waste bits on zeros, and a new layout reclaims them
Breaking the 1.58-bit Barrier for Ternary LLMs
Ternary LLMs store weights as {-1, 0, +1}, but standard five-trit packing assumes all three symbols are equally likely and costs up to 1.625 bits per weight. Measuring 29 ternary models reveals zeros make up as much as 51.5% of weights. BITCOS, a distribution-adaptive layout using a presence bitmap and compacted sign vector, stores weights more compactly in 26 of 29 models, reaching 1.485 bits per weight and up to 1.28× faster matrix-vector multiplication, with decode throughput gains up to 1.18× on CPUs and 1.27× on GPUs.
We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to 51.5% of all weights.
- c7b
> We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to 51.5% of all weights. Motivated by this finding, we introduce BITCOS, a simple distribution-adaptive layout
I honestly assumed that's how they already work. I have to admit that I even explained it like that to a friend. Why on earth wouldn't you design it like that from the start (talking about the adaptive, not the measure part; just sacrifice a few bits to clarify your encoding and save a ton of bits)?
- CodesInChaos
I'm surprised that a variable length encoding like this is usable directly as in memory format and not just as storage/transfer format.
- infogulch
So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.
If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
- om8
Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
- yalok
sounds like a perfect fit for ASIC-optimized models (where matrix ops could be supported directly in BITCOS format, potentially) & achieving record power efficiency for on-device inference.
And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done...
0. https://arxiv.org/pdf/2402.17764 - The Era of 1-bit LLMs:
All Large Language Models are in 1.58 Bits