Repository navigation
Replies: 9 comments 4 replies
|
We are really looking forward to this support. |
|
found this one. |
|
Gemma MTP Drafter is pretty useful, I hope this support could be merged as soon as possible. |
|
I have a fork with working Gemma 4 MTP speculative decoding (tested on E4B, ~70-87% acceptance depending on content, up to ~60% throughput improvement). Not PR-quality but functional, linking for reference: https://github.com/reffdev/llama.cpp/tree/gemma4-mtp |
|
Is this the recommended approach? Still confused on MTP vs Spec decoding. I saw this detailed post of using speculative decoding but I'm unable to reproduce the numbers with some quick Q&A. As with Llama 3.3, etc. speculative decoding is actually slower. Both draft and main model fit in VRAM Ubuntu 24.04 + RADV Vulkan
|
|
Some numbers from a grammar-constrained workload, since most reports here are code or prose. Setup: Gemma 4 E2B, Q4_0 QAT weights (Ollama's Acceptance from the server's
One caution when reading acceptance figures: a run that fell into a repetition loop at temperature 0 reported 91.7 % cumulative acceptance, because the drafter predicts repeated text perfectly. Check output length before trusting a high rate. All timings are wall clock per output token and include a roughly 2k-token prefill. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
I am trying to use the newly released Google Gemma 4 MTP Drafter models (Assistant models) with llama.cpp for speculative decoding. Specifically, I am testing gemma-4-26B-A4B-it-assistant.
Currently, the convert_hf_to_gguf.py script does not recognize the architecture Gemma4AssistantForCausalLM.
Describe the solution you'd like
Support for converting and running Gemma 4 Assistant models in GGUF format. These models are crucial for enabling speculative decoding (using the --draft flag) to speed up inference for the larger Gemma 4 26B/31B models.
Describe alternatives you've considered
I attempted to manually bypass the architecture check by modifying config.json to gemma2, but the conversion failed with the following error due to missing tensor mappings:
ValueError: Can not map tensor 'model.layers.0.layer_scalar'
It seems these models introduced new scaling tensors (like layer_scalar) that are not yet mapped in the current Gemma2Model or GemmaModel classes in convert_hf_to_gguf.py.
Additional context
Model Link: google/gemma-4-26B-A4B-it-assistant
Environment: Windows 10, CUDA 13, Python 3.11
Goal: Use these models as drafters to achieve higher tokens-per-second on consumer hardware (RTX 5060 Ti).
Since Gemma 4 is a major release, having official support for its dedicated assistant models would significantly benefit the community.
All reactions