Started April 19, 2026
I wanted to actually understand quantization — not just read about it. So I grabbed llama.cpp, pulled Qwen-3-8B off HuggingFace, and ran the same prompts through Q4_K_M and Q2_K to see what breaks.
Inspired by Nunchaku (which does something similar but for diffusion models). This is for LLMs.
@douzog 🥷👾
Three quantization levels, same model, same prompts:
| Model | Bits | Size | Speed |
|---|---|---|---|
| Qwen-3-8B F16 (full) | 16-bit | ~16 GB | Slowest |
| Qwen-3-8B Q4_K_M | 4-bit | ~5 GB | Fast |
| Qwen-3-8B Q2_K | 2-bit | ~3 GB | Fastest |
For each run I track speed (tok/s), latency, RAM, and what the model actually says.
Qwen/Qwen-3-8B— HuggingFace base model ID used for the benchmark.qwen3-8b-f16.gguf— full FP16 GGUF copy of Qwen-3-8B.qwen3-8b-q4km.gguf— Q4_K_M quantized variant.qwen3-8b-q2k.gguf— Q2_K quantized variant.
Q2_K loads faster and uses way less RAM (~0.36 GB vs ~0.89 GB), but the responses are noticeably worse on anything that requires reasoning. Q4_K_M is the sweet spot.
Nunchaku targets diffusion models (FLUX, SANA). This is for text generation — different tools for different jobs.
You'll need: Python 3.10+, cmake, ~20 GB disk space.
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
# Apple Silicon
cmake -B build -DGGML_METAL=ON
cmake --build build --config Release -j$(nproc)
# NVIDIA GPU
# cmake -B build -DGGML_CUDA=ON
# cmake --build build --config Release -j$(nproc)hf download Qwen/Qwen-3-8B --local-dir ./models/qwen3-8b
pip install -r requirements.txt
python convert_hf_to_gguf.py models/qwen3-8b --outfile ../qwen3-8b-f16.gguf
./build/bin/llama-quantize ../qwen3-8b-f16.gguf ../qwen3-8b-q4km.gguf Q4_K_M
./build/bin/llama-quantize ../qwen3-8b-f16.gguf ../qwen3-8b-q2k.gguf Q2_KYou should end up with three files:
qwen3-8b-f16.gguf (~16 GB)
qwen3-8b-q4km.gguf (~5 GB)
qwen3-8b-q2k.gguf (~3 GB)
pip install llama-cpp-python psutil matplotlib numpy
jupyter notebook qwen_benchmark.ipynbWeights come out of training as FP32 (or FP16). Quantization maps that float distribution down to a lower-bitwidth integer:
x_q = round(x / scale) + zero_point
x_dequant = (x_q - zero_point) * scale
scale and zero_point are computed per-block. The gap between x and x_dequant is your quantization error — you're trying to minimize it across the whole weight distribution.
Q 4 _K_M
│ │ │ │
│ │ │ └── M = Medium (balanced across layers)
│ │ └──── K = K-quant (smarter block quantization)
│ └──────── 4 = 4 bits per weight
└─────────── Q = Quantized
The older naive quants (Q4_0, Q4_1) used 32-weight blocks with one scale per block. K-quants use 256-weight blocks and a two-level scale — a high-precision super-block scale (FP16) with sub-block scales relative to it. Better range coverage, less error.
The _M suffix means mixed precision: most layers run at 4-bit, but attention weights and other precision-sensitive layers stay at 6-bit. Q4_K_M is effectively ~4.5-bit on average. That's where most of the quality comes from.
At 2-bit you have 4 representable values per weight. That's not enough for:
- Attention projections — QKV layers are precision-sensitive; small errors compound across heads
- MLP nonlinearities — quantization noise gets amplified
- Deep reasoning — errors stack across 32+ layers and you drift far from the intended representation
Q4 hits a sweet spot because transformer weight distributions are roughly Gaussian and tight — 16 buckets covers that well. 4 doesn't.
llama.cpp does PTQ (post-training quantization) — quantize a trained model after the fact, no retraining. Fast to produce, but leaves quality on the table.
QAT (quantization-aware training) fakes quantization during the forward pass so the model learns to tolerate it. Much better at low bitwidths — this is closer to what Nunchaku does for diffusion models.
For LLMs, GPTQ and AWQ are smarter PTQ alternatives: they use calibration data to minimize layer-wise reconstruction error instead of just rounding. Both beat K-quants at equivalent bitwidths, at the cost of more compute during quantization.
The difference between Q2 and Q4 is just bits per weight:
8B params × 2 bits / 8 = ~2 GB (Q2_K)
8B params × 4.5 bits / 8 = ~4.5 GB (Q4_K_M effective)
The smaller delta in the benchmark (~0.36 vs ~0.89 GB) is because llama.cpp pages layers depending on your --n-gpu-layers setting — it doesn't load the whole model at once.
The Q2_K tradeoff in practice:
Q4_K_M → "The mitochondria is the powerhouse of the cell because..."
Q2_K → "The mitochondria powerhouse cell energy ATP..."
Qwentify/
├── README.md
├── qwen_benchmark.ipynb
├── qwen_benchmark.png
└── qwen_benchmark_results.csv
| Problem | Fix |
|---|---|
no such file: ./build/bin/llama-quantize |
cmake build didn't finish — rerun those steps |
cmake not found |
brew install cmake on Mac, sudo apt install cmake -y on Linux |
| OOM during quantization | F16 needs ~18 GB RAM. For 35B you're looking at ~80 GB — just grab a pre-quantized GGUF from HuggingFace |
