bzip3: A spiritual successor to BZip2 that compresses 26 years of Perl source code to 546 MB

bzip3: A spiritual successor to BZip2 that compresses 26 years of Perl source code to 546 MB

bzip3 is a new compression tool and library that aims to be a better, faster, and stronger successor to the classic BZip2. It combines a Burrows-Wheeler transform (BWT) with a context-mixing entropy coder and a Lempel-Ziv+Prediction (LZP) pass, achieving significantly higher compression ratios, especially on text and code. In a benchmark compressing every version of Perl 5 ever released (262 tarballs), bzip3 with a 511 MB block size produced a 546 MB archive, outperforming xz, bzip2, and Zstandard in both size and decompression speed. The project is actively maintained and available under the LGPLv3 license.

Bzip3's performance is heavily dependent on the compiler. x64 Linux clang13 builds usually can go as high as 17MiB/s compression and 23MiB/s decompression per thread.
  1. altairprime

    Previously:

    “Hi, tool author here.” A useful explanation of Burrows-Wheelers transform as used by bzip3: https://news.ycombinator.com/item?id=42902407

    “bzip3 is not yet listed on the large text compression benchmark” It is now: https://mattmahoney.net/dc/text.html

    (2 years ago, 176 comments) https://news.ycombinator.com/item?id=42899713

    (4 years ago, 104 comments) https://news.ycombinator.com/item?id=31324439

  2. 8organicbits

    I was processing compressed .jsonl files recently (JSON lines format). I found that lzma gave a much better compression than gzip or bzip2, which helps for archival costs, but it's challenging to work with as software support is lacking. I do duckdb processing which supports gzip transparently. There's an extension for bzip2, but not for lzma or bzip3.

    I ended up using gzip because it's best supported by the software I use and most likely to have support in software I adopt. But it gave the worst compression results of the options I tried. These bzip3 numbers certainly give me FOMO...

  3. ot

    The benchmarks are disingenuous, to the point of looking cherry-picked. The block size for bzip3 is set to 512MB, but the window size for zstd is left to its default (8MB I believe for high levels). So in this corpus, which is made up of all versions of Perl source code concatenated, the window is too small to see all the identical files and just match them.

    Also corpora made out of very long repetitions are pretty much the best case scenario for BWT-based compressors.

    If we match the window size of zstd to that of bzip3 we get dramatically different results:

    % gzcat *.gz | time zstd -T8 -16 | wc -c # baseline

    2819113884

    zstd -T8 -16 2054.50s user 3.47s system 783% cpu 4:22.80 total

    % gzcat *.gz | time zstd -T8 -16 --long=29 | wc -c

    196405076

    zstd -T8 -16 --long=29 1083.06s user 2.41s system 783% cpu 2:18.55 total

    Almost 15x smaller than the baseline, and more than 2x smaller than bzip3, also CPU time halves (since long matches are found earlier, so there's less work to do).

    (the baseline number is slightly different because I don't have the exact Perl version set used by the author)

    Also, in the benchmarks using lrzip, which would make the window size less relevant, zstd is not even compared.

  4. JdeBP

    An interesting unintentional benchmark is to go to https://github.com/iczelia/bzip3/releases and see to what degree bzip3 compresses its own release archives; and go to https://github.com/iczelia/bzip3/blob/master/.github/workflo... to see what options have been chosen for the other compressors here.

  5. CodesInChaos

    I think this should include benchmarks zstd with larger windows and long range mode. I wouldn't be surprised if the window they used is smaller than an individual version tar file, which would prevent useful compression.

  6. amelius

    > However, the complexity of the algorithms, and, in particular, the presence of various special cases in the code which occur with very low but non-zero probability make it impossible to rule out the possibility of bugs remaining in the program.

    Sounds like perhaps a nice testcase for formalization + AI?

  7. BorisMelnik

    faster than gzip now or still slower? (sorry I did not read readme)

  8. whatever1

    I think compression algorithms are ripe for significant improvement with LLMs. It’s an ideal candidate problem you can have in a closed loop evaluation, and you can just let agent try things.

More from this day

2026-09-07