Skip to content
imazenPublic

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Repository files navigation

zenflate CI crates.io lib.rs docs.rs MSRV license

Pure Rust DEFLATE / zlib / gzip. Compression spans effort levels 0–200 across seven strategies (and can emit byte-identical output to C libdeflate on demand), with whole-buffer and streaming decompression plus SIMD Adler-32 / CRC-32. #![forbid(unsafe_code)] by default (with an opt-in unchecked fast path) and no_std-friendly: compression and streaming decompression require alloc, while whole-buffer decompression works without alloc. With std, the fixed-Huffman table cache allocates once per process.

Quick start

[dependencies]
zenflate = "0.4"
use zenflate::{Compressor, Decompressor, CompressionLevel, Unstoppable};

let data = b"the quick brown fox jumps over the lazy dog, again and again";

// Compress with the balanced preset (raw DEFLATE; zlib_/gzip_ variants share this shape).
let mut compressor = Compressor::new(CompressionLevel::balanced());
let mut packed = vec![0u8; Compressor::deflate_compress_bound(data.len())];
let n = compressor.deflate_compress(data, &mut packed, Unstoppable).unwrap();

// Decompress into a caller-sized buffer — its length is your hard size cap.
let mut out = vec![0u8; data.len()];
let r = Decompressor::new()
    .deflate_decompress(&packed[..n], &mut out, Unstoppable)
    .unwrap();
assert_eq!(&out[..r.output_written], data);

Need gzip/zlib framing, streaming, parallel gzip, cancellation, or fine-grained effort control? Those are covered below.

Usage

Compress

use zenflate::{Compressor, CompressionLevel, Unstoppable};

let data = b"Hello, World! Hello, World! Hello, World!";
let mut compressor = Compressor::new(CompressionLevel::balanced());

let bound = Compressor::deflate_compress_bound(data.len());
let mut compressed = vec![0u8; bound];
let compressed_len = compressor
    .deflate_compress(data, &mut compressed, Unstoppable)
    .unwrap();
let compressed = &compressed[..compressed_len];

Decompress

use zenflate::{Decompressor, Unstoppable};

let mut decompressor = Decompressor::new();
let mut output = vec![0u8; original_len];
let result = decompressor
    .deflate_decompress(compressed, &mut output, Unstoppable)
    .unwrap();
// result.input_consumed — bytes of compressed data consumed
// result.output_written — bytes of decompressed data produced

For gzip and zlib, use gzip_decompress / zlib_decompress (identical shape).

The one-shot decoders may overwrite up to about 32 bytes of output past output_written when the buffer is larger than the decoded data (the fast loop copies matches in fixed-size chunks). Bytes past output_written are not preserved, so don't decode into part of a buffer whose tail you need.

Server safety — bound the output. The one-shot decompressors write into the &mut [u8] you pass, so that buffer is the size cap: for untrusted input you don't know the decompressed length up front (the gzip trailer is attacker- controlled), so size output to your maximum and decompression returns an error rather than over-allocating. If you instead use the streaming StreamDecompressor (which grows its own buffer), cap it explicitly with .with_max_output_size(Some(max_bytes)) — otherwise a small "zip bomb" can expand without bound.

// gzip into a hard-capped buffer (rejects anything larger):
let mut out = vec![0u8; 100 * 1024 * 1024]; // 100 MiB ceiling
match Decompressor::new().gzip_decompress(gzip_bytes, &mut out, Unstoppable) {
    Ok(r) => { /* r.output_written bytes are valid */ }
    Err(e) => { /* malformed input or output exceeds the 100 MiB ceiling */ }
}

Streaming decompression

For inputs that don't fit in memory or arrive incrementally. Works with &[u8] (zero overhead) or any std::io::BufRead via BufReadSource.

Construct with deflate/zlib/gzip (each takes the source plus an output buffer capacity — DEFAULT_CAPACITY is 64 KiB), then drive the fill → peek → advance loop until is_done():

use zenflate::{StreamDecompressor, DEFAULT_CAPACITY};

// From a slice (`&[u8]` is a zero-overhead source):
let mut stream = StreamDecompressor::deflate(compressed_data, DEFAULT_CAPACITY);
while !stream.is_done() {
    stream.fill()?;             // pull from source, decompress into the buffer
    let chunk = stream.peek();  // borrow the available decompressed output
    // process chunk...
    let n = chunk.len();
    stream.advance(n);          // mark consumed, freeing buffer space
}

// From a BufRead (std only):
use zenflate::BufReadSource;
let file = std::io::BufReader::new(std::fs::File::open("data.gz").unwrap());
let mut stream = StreamDecompressor::gzip(BufReadSource::new(file), DEFAULT_CAPACITY);
// stream also implements Read + BufRead

Untrusted input / decompression bombs. The whole-buffer Decompressor is naturally bounded by the output slice you pass it. The streaming API produces output incrementally, so for untrusted data cap the total with with_max_output_size (decoding then errors instead of allocating past the cap); a stall guard also rejects streams that emit thousands of empty blocks without progress:

let mut stream = StreamDecompressor::gzip(compressed_data, DEFAULT_CAPACITY)
    .with_max_output_size(Some(64 * 1024 * 1024)); // DecompressionError::OutputLimitExceeded past 64 MiB

Formats

All three DEFLATE-based formats are supported:

// Raw DEFLATE
compressor.deflate_compress(data, &mut out, Unstoppable)?;
decompressor.deflate_decompress(compressed, &mut out, Unstoppable)?;

// zlib (2-byte header + DEFLATE + Adler-32)
compressor.zlib_compress(data, &mut out, Unstoppable)?;
decompressor.zlib_decompress(compressed, &mut out, Unstoppable)?;

// gzip (10-byte header + DEFLATE + CRC-32)
compressor.gzip_compress(data, &mut out, Unstoppable)?;
decompressor.gzip_decompress(compressed, &mut out, Unstoppable)?;

Compression levels

Pick a preset or dial in a specific effort from 0 to 200:

use zenflate::CompressionLevel;

// Named presets
CompressionLevel::none()      // effort 0  — store (no compression)
CompressionLevel::fastest()   // effort 1  — turbo hash table
CompressionLevel::fast()      // effort 10 — greedy hash chains
CompressionLevel::balanced()  // effort 15 — lazy matching (default)
CompressionLevel::high()      // effort 22 — double-lazy matching
CompressionLevel::best()      // effort 30 — near-optimal parsing

// Fine-grained control (0-200, clamped)
CompressionLevel::new(12)     // lazy matching, mid-range
CompressionLevel::new(25)     // near-optimal, fast end
CompressionLevel::new(46)     // Zopfli-style full-optimal, 30 iterations

// Byte-identical C libdeflate compatibility (0-12)
CompressionLevel::libdeflate(6)
Preset Effort Strategy Description
none() 0 Store Framing only, no compression
fastest() 1 Turbo Maximum throughput
fast() 10 Greedy Hash chains — big ratio jump over turbo
balanced() 15 Lazy Lazy matching — good default
high() 22 Lazy2 Double-lazy — best before near-optimal
best() 30 Near-optimal Best compression ratio

Effort levels map to seven strategies:

Effort Strategy Notes
0 Store No compression
1-4 Turbo Single-entry hash table, fastest
5-9 FastHt 2-entry hash table, increasing match length
10 Greedy Hash chains with greedy matching
11-17 Lazy Hash chains with lazy matching
18-22 Lazy2 Double-lazy matching
23-30 Near-optimal Near-optimal parsing via binary trees
31-200 FullOptimal Zopfli-style iterative optimal parsing (iterations = effort − 16); very slow, maximum density

Higher effort within a strategy increases search depth and match quality. Strategy transitions (e.g. e9→e10, e10→e11) can occasionally produce slightly larger output on specific inputs due to algorithmic differences. Use CompressionLevel::monotonicity_fallback() to detect and handle these transitions — it returns the previous strategy's max effort so you can compare both and pick the smaller result.

Reuse Compressor and Decompressor across calls to avoid re-initialization.

Recommended effort levels

For most uses, balanced() (effort 15) is a good default. Use fast() (effort 10) when speed matters more than the last few percent of compression.

PNG image data

CompressionLevel::png(effort) is tuned for PNG IDAT streams: filtered scanlines with long byte runs and literal-heavy residuals.

Effort Encoder
png(1) Ultra-fast: literals and zero runs, one Huffman table per stream
png(2) Runs only, exact Huffman tables per block
png(3) Runs + hashed repeats of 8+ bytes
png(4..=9) Runs + hashed repeats of 5+ bytes, hash chains of growing depth
png(10..=18) Lazy matching, search depth 16 → 800
png(19..=22) Near-optimal parsing, search depth 16 → 35
png(23..=30) new(23..=30)'s near-optimal settings
png(31..) Same as new(31..) (full optimal parsing)

From png(3) through png(30) a runs-only guard compares the selected blocks against a runs-only parse and writes the cheaper Huffman encoding. After a clear loss, the guard skips the next three blocks. This improves flat-colour art but does not guarantee output no larger than png(2). From png(10) through png(26) target block boundaries come from the input alone; near-optimal parsing can still end a block early if its match cache fills. png(27..=30) can also split inside those input-derived segments (about 0.1% smaller on the measured set, less strictly nested). monotonicity_fallback() names the lower level to compare against at each change of algorithm.

use zenflate::{Compressor, CompressionLevel, Unstoppable};

let mut compressor = Compressor::new(CompressionLevel::png(6));
let mut idat = vec![0u8; Compressor::zlib_compress_bound(filtered_rows.len())];
let size = compressor.zlib_compress(&filtered_rows, &mut idat, Unstoppable)?;

Measured on 86 PNG filtered streams (64–1024 px), Ampere Altra Neoverse-N1, one core, each library compressing the same bytes:

zenflate Ratio MB/s Nearby levels of other libraries
png(1) 3.078 1129 fdeflate ultra-fast: 2.835 @ 707
png(2) 3.360 466
png(3) 3.573 215 libdeflate 1: 3.611 @ 197, zlib-rs 1: 2.590 @ 207
png(4) 3.651 171 miniz_oxide 1: 3.199 @ 168
png(9) 3.771 98
png(10) 3.823 60 libdeflate 6: 3.839 @ 65
png(12) 3.903 35 zlib-rs 6: 3.863 @ 49
png(16) 3.948 18 miniz_oxide 6: 3.849 @ 21, libdeflate 9: 3.928 @ 15
png(18) 3.958 14 zlib-rs 9: 3.973 @ 10, miniz_oxide 9: 3.911 @ 8
png(19) 4.087 8
png(23) 4.115 7
png(26) 4.135 4
png(28) 4.143 3
png(30) 4.145 2 libdeflate 12: 4.144 @ 2

Higher levels can still produce a slightly larger file than a lower one on some images: on this set by at most 0.81% from png(19) through png(26). libdeflate 6 beats png(10) on both size and speed; png(30) is smaller than libdeflate 12 and png(28) faster. Per-image data: benchmarks/png_ladder_ramp_2026-10-07.txt, benchmarks/png_ladder_final_2026-10-07.txt; held-out validation of png(1..=9): benchmarks/png_mode_2026-10-06.md.

PNG strips for parallel encode and decode

zenflate::png::{StripCompressor, StripDecoder} split one PNG zlib stream into independent strips of whole rows (PNG's iDOT layout). Each strip is compressed without history from earlier strips, so strips can be compressed on separate threads, and the concatenation is still one valid zlib stream that any decoder reads. StripDecoder inflates one strip on its own (streaming, fill/peek/advance) and reports whether it ended where the next strip begins, so an iDOT-aware decoder can inflate strips in parallel and verify the result against the stream's Adler-32.

use zenflate::png::StripCompressor;
use zenflate::{CompressionLevel, Unstoppable, adler32, adler32_combine};

let mut c = StripCompressor::new(CompressionLevel::png(4));
let mut z = c.zlib_header().to_vec();
let mut adler = 1;
for (k, strip) in strips.iter().enumerate() {
    let mut out = vec![0u8; StripCompressor::bound(strip.len())];
    let n = c.compress(strip, k + 1 == strips.len(), &mut out, Unstoppable)?;
    z.extend_from_slice(&out[..n]);
    adler = adler32_combine(adler, adler32(1, strip), strip.len());
}
z.extend_from_slice(&adler.to_be_bytes());

Parallel gzip compression

use zenflate::{Compressor, CompressionLevel, Unstoppable};

let mut compressor = Compressor::new(CompressionLevel::balanced());
let bound = Compressor::gzip_compress_bound(data.len()) + num_threads * 5;
let mut compressed = vec![0u8; bound];
let size = compressor
    .gzip_compress_parallel(data, &mut compressed, 4, Unstoppable)
    .unwrap();

Splits input into chunks with 32KB dictionary overlap, compresses in parallel, concatenates into a valid gzip stream. Scaling depends on input size and content.

Cancellation

All compression and whole-buffer decompression methods accept a stop parameter implementing the Stop trait. Pass Unstoppable to disable cancellation, or implement Stop to check a flag periodically:

use zenflate::{Stop, StopReason, Unstoppable};

// Unstoppable — never cancels
compressor.deflate_compress(data, &mut out, Unstoppable)?;

// Custom cancellation
struct MyStop { cancelled: std::sync::Arc<std::sync::atomic::AtomicBool> }
impl Stop for MyStop {
    fn check(&self) -> Result<(), StopReason> {
        if self.cancelled.load(std::sync::atomic::Ordering::Relaxed) {
            Err(StopReason)
        } else {
            Ok(())
        }
    }
}

Streaming decompression doesn't take a Stop parameter — the caller controls the loop and can stop between fill() calls.

Features

Feature Default Effect
std yes std::io::{Read, BufRead} integration (BufReadSource)
alloc yes (via std) Streaming decompression
compress yes Compressor / CompressionLevel (implies alloc)
simd yes Runtime-dispatched SIMD checksums and matchfinder multiversioning (via archmage); without it, scalar paths
avx512 yes AVX-512 SIMD tiers (implies simd)
threads yes Parallel gzip (gzip_compress_parallel, implies compress); disable for thread-less wasm32
unchecked no Elide bounds checks in compression hot paths (+0-12% compression speed)

Decompression works in no_std without alloc; all state is stack-allocated.

For a minimal, fast-to-compile decoder, disable default features:

zenflate = { version = "0.4.0", default-features = false, features = ["std"] }

That decode-only configuration has a single direct dependency (enough) — no proc macros, no SIMD — and still decodes all three formats with checksum verification (scalar Adler-32/CRC-32).

Migrating from 0.3: with default-features = false, add compress if you compress and simd if you want SIMD checksums; both were previously implied by alloc / always-on.

Performance

Measured on zenflate 0.4.0, AMD Ryzen 9 7950X (Zen 4), Linux/WSL2, safe mode (the default — forbid(unsafe_code), no unchecked), no -C target-cpu=native (runtime SIMD dispatch only). The full head-to-head — x86 + aarch64, the whole Rust ecosystem, real corpora, per-host — is committed at benchmarks/deflate_rust_ecosystem_2026-07-13.md. One machine, one run set — re-measure before quoting externally.

Compression (1 MB mixed synthetic, median of n=100, lower is better):

Library L1 L6 L12 / max
zenflate 5.53 ms 6.08 ms 8.31 ms
libdeflate (C) 4.95 ms 6.03 ms 17.02 ms
zlib-rs 5.35 ms 12.69 ms 14.56 ms (L9)
miniz_oxide 2.81 ms 14.63 ms 15.33 ms (L9)

At L6 zenflate matches C and is ~2× faster than every other Rust crate; at L12 it is ~2× faster than C (a different, faster near-optimal algorithm). Level numbers are not equivalent across libraries — compare at matched ratio (see the benchmark file). Via CompressionLevel::libdeflate(n), zenflate emits byte-identical output to C libdeflate at every level.

Decompression (1,000,000 bytes, compressed at zenflate L6, lower is better):

Data zenflate libdeflate (C) flate2 (zlib-rs) miniz_oxide
Sequential 45.9 µs 35.2 µs 37.9 µs 89.0 µs
Mixed 1.31 ms 1.24 ms 1.54 ms 1.81 ms
Photo 1.51 ms 1.44 ms 1.73 ms 2.10 ms

These are the times from the committed ecosystem record linked above. On its synthetic mixed/photo inputs zenflate was the fastest Rust decoder measured. The same record's Silesia results vary by file and do not support a universal lead.

PNG streams (106 IDAT streams, 64–2560 px, median time per image relative to fdeflate, measured 2026-10-07 after 0.4.0, Core Ultra 7 265K one core): one-shot 0.96×, streaming 0.98×; Ryzen 9 9950X3D (Zen 5) one-shot 0.98× with the AVX-512 build; Neoverse-N1 one-shot 0.83×, streaming 0.87× (benchmarks/chunk32_copy_2026-10-07.txt, benchmarks/oneshot_v4_2026-10-07.txt).

Checksums (1 MiB, median of five interleaved rounds on the same Zen 4 host, 2026-07-13):

Algorithm Without avx512 With avx512
Adler-32 77.4 GiB/s 112.8 GiB/s
CRC-32 18.4 GiB/s 78.2 GiB/s

Source: AVX-512 checksum A/B.

How it works

zenflate started as a port of Eric Biggers' libdeflate and has grown into its own implementation. The core decompressor, matchfinders, Huffman construction, and block splitting trace back to libdeflate. On top of that foundation, zenflate pulls in techniques from several other projects and adds original work:

  • Effort-based compression (0-200) with seven strategies and named presets, replacing libdeflate's fixed 0-12 levels. Includes two original matchfinder designs (turbo, fast HT) for the low-effort range.
  • Full-optimal compression (Zopfli-style iterative squeeze), ported from zenzop with Katajainen bounded package-merge for optimal length-limited Huffman codes.
  • Multi-strategy Huffman optimization combining Brotli-inspired frequency smoothing, Zopfli-style RLE optimization, and max-bits sweeps to find the smallest encoding per block.
  • Parallel gzip compression using pigz-style chunking with 32KB dictionary overlap and combined CRC-32 via GF(2) matrix.
  • Streaming decompression via a pull-based API that works in no_std + alloc.
  • Snapshot/restore (CompressorSnapshot) for branching compression state — try different inputs from the same point and pick the best result (designed for PNG filter selection).
  • Cancellation via the Stop trait for cooperative interruption.

Safe Rust throughout (#![forbid(unsafe_code)] by default), with an opt-in unchecked feature for bounds-check elimination in compression hot paths. SIMD acceleration for checksums (AVX2/AVX-512/PCLMULQDQ on x86, NEON/PMULL on aarch64, simd128 on WASM) via archmage with zero unsafe.

zenflate can produce byte-identical output to libdeflate at every level (via CompressionLevel::libdeflate(n)), and runs at roughly 0.8-0.9x the speed of the C original depending on level and data. The gap comes from register pressure differences and bounds checking.

Acknowledgments

  • libdeflate by Eric Biggers — decompressor, matchfinders (hash table, hash chains, binary trees), Huffman construction, block splitting, near-optimal parser, checksum implementations
  • Zopfli by Lode Vandevenne and Jyrki Rissanen (Google) — full-optimal parsing concept, iterative cost refinement, optimize_huffman_for_rle (Zopfli-style variant)
  • zenzop — Rust Zopfli port used as the source for katajainen, squeeze, and block splitter modules
  • Brotli (Google) — frequency smoothing algorithm for Huffman RLE encoding
  • pigz by Mark Adler — parallel gzip chunking strategy with dictionary overlap
  • fdeflate (image-rs) — the PNG ultra-fast, runs-only and greedy compressors that png(1..=9) adapt, and the double-literal decode tables and chunked match copy in the inflate loop

What's different from libdeflate

CompressionLevel::libdeflate(n) produces byte-identical output to C. The recommended effort-based API (CompressionLevel::new(n)) uses different algorithms and tuning at every level:

Effort Strategy Matchfinder Encoding vs libdeflate
0 Store — — Same
1-4 Turbo Single-entry hash, limited skip updates Standard Original matchfinder, not in libdeflate
5-9 FastHt 2-entry hash, limited skip updates Standard Original matchfinder, not in libdeflate
10 Greedy Hash chains Standard good_match early-exit (libdeflate: disabled)
11-17 Lazy Hash chains Standard good_match/max_lazy tuning curves (libdeflate: disabled)
18-22 Lazy2 Hash chains Standard good_match/max_lazy tuning (libdeflate: disabled)
23-25 NearOptimal Binary trees Exhaustive precode search Multi-strategy precode flag search
26-27 NearOptimal Binary trees + multi-strategy Huffman + Brotli/Zopfli RLE smoothing, reduced max_bits sweep
28-30 NearOptimal Binary trees + diversified optimization + randomized cost model, 20-30 passes (libdeflate: 2-10)
31+ FullOptimal Zopfli hash chains Katajainen package-merge Entirely different algorithm (from zenzop)

At effort 10-22, the core matching algorithms are the same as libdeflate (greedy, lazy, double-lazy with hash chains), but zenflate adds good_match and max_lazy early-exit thresholds that libdeflate leaves disabled. These let the compressor skip deep chain searches and lazy evaluations when it already has a good enough match, trading a small amount of compression ratio for speed at lower effort levels.

At effort 23+, the near-optimal parser is the same backward DP as libdeflate, but the block encoding pipeline diverges: multi-strategy Huffman code construction tries Brotli-inspired and Zopfli-style frequency smoothing with max-bits sweeps to find smaller encodings. At effort 28+, the optimizer runs 20-30 passes with randomized cost diversification instead of libdeflate's fixed 2-10 passes.

MSRV

The minimum supported Rust version is 1.89.

AI-Generated Code Notice

Developed with Claude (Anthropic). Not all code manually reviewed. Review critical paths before production use.

License

Dual-licensed: AGPL-3.0 or commercial.

I've maintained and developed open-source image server software — and the 40+ library ecosystem it depends on — full-time since 2011. Fifteen years of continual maintenance, backwards compatibility, support, and the (very rare) security patch. That kind of stability requires sustainable funding, and dual-licensing is how we make it work without venture capital or rug-pulls. Support sustainable and secure software; swap patch tuesday for patch leap-year.

Our open-source products

Your options:

  • Startup license — $1 if your company has under $1M revenue and fewer than 5 employees. Get a key →
  • Commercial subscription — Governed by the Imazen Site-wide Subscription License v1.1 or later. Apache 2.0-like terms, no source-sharing requirement. Sliding scale by company size. Pricing & 60-day free trial →
  • AGPL v3 — Free and open. Share your source if you distribute.

See LICENSE-COMMERCIAL for details.

Upstream code from ebiggers/libdeflate is licensed under MIT. Our additions and improvements are dual-licensed (AGPL-3.0 or commercial) as above.

Image tech I maintain

Codecs ¹ zenjpeg · zenpng · zenwebp · zengif · zenavif · zenjxl · zenjxl-decoder · jxl-encoder · zenbitmaps · heic · zentiff · zenpdf · zensvg · zenjp2 · zenraw · ultrahdr
Codec internals zenrav1e · rav1d-safe · zenravif · zenavif-parse · zenavif-serialize
Compression zenflate · zenzop · zenzstd
Processing zenresize · zenquant · zenblend · zenfilters · zensally · zentone
Pixels & color zenpixels · zenpixels-convert · linear-srgb · garb · zenyuv
Pipeline & framework zenpipe · zencodec · zencodecs · zenlayout · zennode · zenwasm · zentract
Metrics zensim · fast-ssim2 · butteraugli · zenmetrics · resamplescope-rs
Pickers & ML zenanalyze · zenpredict · zenpicker · zenanalyze-api
Test corpora codec-corpus · imazen-26
Products Imageflow image engine (.NET · Node · Go) · Imageflow Server · ImageResizer (C#)

¹ pure-Rust, #![forbid(unsafe_code)] codecs, as of 2026

General Rust awesomeness

zenbench · archmage · magetypes · enough · whereat · cargo-copter · zenutils

Open source · @imazen · @lilith · lib.rs/~lilith

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages