Ternary LLMs waste bits on zeros, and a new layout reclaims them

Breaking the 1.58-bit Barrier for Ternary LLMs

Ternary LLMs store weights as {-1, 0, +1}, but standard five-trit packing assumes all three symbols are equally likely and costs up to 1.625 bits per weight. Measuring 29 ternary models reveals zeros make up as much as 51.5% of weights. BITCOS, a distribution-adaptive layout using a presence bitmap and compacted sign vector, stores weights more compactly in 26 of 29 models, reaching 1.485 bits per weight and up to 1.28× faster matrix-vector multiplication, with decode throughput gains up to 1.18× on CPUs and 1.27× on GPUs.

We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to 51.5% of all weights.
  1. c7b

    > We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to 51.5% of all weights. Motivated by this finding, we introduce BITCOS, a simple distribution-adaptive layout

    I honestly assumed that's how they already work. I have to admit that I even explained it like that to a friend. Why on earth wouldn't you design it like that from the start (talking about the adaptive, not the measure part; just sacrifice a few bits to clarify your encoding and save a ton of bits)?

  2. CodesInChaos

    I'm surprised that a variable length encoding like this is usable directly as in memory format and not just as storage/transfer format.

  3. infogulch

    So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

    If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.

  4. om8

    Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.

  5. yalok

    sounds like a perfect fit for ASIC-optimized models (where matrix ops could be supported directly in BITCOS format, potentially) & achieving record power efficiency for on-device inference.

    And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done...

    0. https://arxiv.org/pdf/2402.17764 - The Era of 1-bit LLMs:

    All Large Language Models are in 1.58 Bits

More from this day

2026-09-16