Skip to content
douzogPublic

About

A tutorial on quantizing llama.cpp with Qwen2-8b (or Qwen3.6-35b if it doesn't break)

Resources

Stars

3 stars

Watchers

0 watching

Forks

Repository files navigation

Qwentify 🥷👾

Started April 19, 2026

I wanted to actually understand quantization — not just read about it. So I grabbed llama.cpp, pulled Qwen-3-8B off HuggingFace, and ran the same prompts through Q4_K_M and Q2_K to see what breaks.

Inspired by Nunchaku (which does something similar but for diffusion models). This is for LLMs.

@douzog 🥷👾


What's in here

Three quantization levels, same model, same prompts:

Model Bits Size Speed
Qwen-3-8B F16 (full) 16-bit ~16 GB Slowest
Qwen-3-8B Q4_K_M 4-bit ~5 GB Fast
Qwen-3-8B Q2_K 2-bit ~3 GB Fastest

For each run I track speed (tok/s), latency, RAM, and what the model actually says.

Model names used in this repo

  • Qwen/Qwen-3-8B — HuggingFace base model ID used for the benchmark.
  • qwen3-8b-f16.gguf — full FP16 GGUF copy of Qwen-3-8B.
  • qwen3-8b-q4km.gguf — Q4_K_M quantized variant.
  • qwen3-8b-q2k.gguf — Q2_K quantized variant.

Results

Qwen Benchmark

Q2_K loads faster and uses way less RAM (~0.36 GB vs ~0.89 GB), but the responses are noticeably worse on anything that requires reasoning. Q4_K_M is the sweet spot.


Why not Nunchaku?

Nunchaku targets diffusion models (FLUX, SANA). This is for text generation — different tools for different jobs.


Setup

You'll need: Python 3.10+, cmake, ~20 GB disk space.

1. Build llama.cpp

git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp

# Apple Silicon
cmake -B build -DGGML_METAL=ON
cmake --build build --config Release -j$(nproc)

# NVIDIA GPU
# cmake -B build -DGGML_CUDA=ON
# cmake --build build --config Release -j$(nproc)

2. Get the model and quantize

hf download Qwen/Qwen-3-8B --local-dir ./models/qwen3-8b

pip install -r requirements.txt
python convert_hf_to_gguf.py models/qwen3-8b --outfile ../qwen3-8b-f16.gguf

./build/bin/llama-quantize ../qwen3-8b-f16.gguf ../qwen3-8b-q4km.gguf Q4_K_M
./build/bin/llama-quantize ../qwen3-8b-f16.gguf ../qwen3-8b-q2k.gguf  Q2_K

You should end up with three files:

qwen3-8b-f16.gguf    (~16 GB)
qwen3-8b-q4km.gguf   (~5 GB)
qwen3-8b-q2k.gguf    (~3 GB)

3. Run the benchmark

pip install llama-cpp-python psutil matplotlib numpy
jupyter notebook qwen_benchmark.ipynb

How quantization actually works

Weights come out of training as FP32 (or FP16). Quantization maps that float distribution down to a lower-bitwidth integer:

x_q       = round(x / scale) + zero_point
x_dequant = (x_q - zero_point) * scale

scale and zero_point are computed per-block. The gap between x and x_dequant is your quantization error — you're trying to minimize it across the whole weight distribution.

What Q4_K_M actually is

Q  4  _K_M
│  │   │ │
│  │   │ └── M = Medium (balanced across layers)
│  │   └──── K = K-quant (smarter block quantization)
│  └──────── 4 = 4 bits per weight
└─────────── Q = Quantized

The older naive quants (Q4_0, Q4_1) used 32-weight blocks with one scale per block. K-quants use 256-weight blocks and a two-level scale — a high-precision super-block scale (FP16) with sub-block scales relative to it. Better range coverage, less error.

The _M suffix means mixed precision: most layers run at 4-bit, but attention weights and other precision-sensitive layers stay at 6-bit. Q4_K_M is effectively ~4.5-bit on average. That's where most of the quality comes from.

Why Q2_K falls apart

At 2-bit you have 4 representable values per weight. That's not enough for:

  • Attention projections — QKV layers are precision-sensitive; small errors compound across heads
  • MLP nonlinearities — quantization noise gets amplified
  • Deep reasoning — errors stack across 32+ layers and you drift far from the intended representation

Q4 hits a sweet spot because transformer weight distributions are roughly Gaussian and tight — 16 buckets covers that well. 4 doesn't.

PTQ vs. QAT (why llama.cpp is fast but imperfect)

llama.cpp does PTQ (post-training quantization) — quantize a trained model after the fact, no retraining. Fast to produce, but leaves quality on the table.

QAT (quantization-aware training) fakes quantization during the forward pass so the model learns to tolerate it. Much better at low bitwidths — this is closer to what Nunchaku does for diffusion models.

For LLMs, GPTQ and AWQ are smarter PTQ alternatives: they use calibration data to minimize layer-wise reconstruction error instead of just rounding. Both beat K-quants at equivalent bitwidths, at the cost of more compute during quantization.

The RAM numbers

The difference between Q2 and Q4 is just bits per weight:

8B params × 2 bits / 8 = ~2 GB   (Q2_K)
8B params × 4.5 bits / 8 = ~4.5 GB  (Q4_K_M effective)

The smaller delta in the benchmark (~0.36 vs ~0.89 GB) is because llama.cpp pages layers depending on your --n-gpu-layers setting — it doesn't load the whole model at once.


The Q2_K tradeoff in practice:

Q4_K_M  →  "The mitochondria is the powerhouse of the cell because..."
Q2_K    →  "The mitochondria powerhouse cell energy ATP..."

Repo

Qwentify/
├── README.md
├── qwen_benchmark.ipynb
├── qwen_benchmark.png
└── qwen_benchmark_results.csv

Troubleshooting

Problem Fix
no such file: ./build/bin/llama-quantize cmake build didn't finish — rerun those steps
cmake not found brew install cmake on Mac, sudo apt install cmake -y on Linux
OOM during quantization F16 needs ~18 GB RAM. For 35B you're looking at ~80 GB — just grab a pre-quantized GGUF from HuggingFace

Links

About

A tutorial on quantizing llama.cpp with Qwen2-8b (or Qwen3.6-35b if it doesn't break)

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages