Skip to content

[Bugfix] Ensure calculated KV scales are applied in attention. - #27232

Merged
ProExpertProg merged 25 commits into
vllm-project:mainfrom
adabeyta:kv_scales_fix_27102
Nov 10, 2025
Merged

ProExpertProg merged 25 commits into
vllm-project:mainfrom
adabeyta:kv_scales_fix_27102

Conversation

@adabeyta

@adabeyta adabeyta commented Oct 21, 2025 •

Copy link
Copy Markdown
Contributor

Purpose:

Resolves bug #27102 .

Test Plan:

Throughput

vllm bench throughput   --model Qwen/Qwen3-8B   --quantization fp8   --kv-cache-dtype fp8_e4m3   --tensor-parallel-size 2   --dataset-name random --input-len 1024 --output-len 256

E2E Correctness
Server with KV scale calculation ON (remove --calculate-kv-scales flag for OFF case):

vllm serve Qwen/Qwen3-8B \
    --tensor-parallel-size 2 \
    --quantization fp8 \
    --kv-cache-dtype fp8_e4m3 \
    --calculate-kv-scales \
    --port 8000

lm_eval

lm_eval --model local-completions --model_args pretrained <model>,base_url=http://0.0.0.0:8000/v1/completions,num_concurrent=50,max_retries=3 --tasks gsm8k --num_fewshot 5 --batch_size auto --limit 100 

Throughput results

Main:calculate_kv_scales=False | enforce_eager=False


Throughput: 1.73 requests/s, 2210.02 total tokens/s, 442.00 output tokens/s
Total num prompt tokens:  1024000
Total num output tokens:  256000



Main:calculate_kv_scales=True | enforce_eager=False

Throughput: 1.73 requests/s, 2211.36 total tokens/s, 442.27 output tokens/s
Total num prompt tokens:  1024000
Total num output tokens:  256000

Main: calulate_kv_scales=True | enforce_eager=True

Throughput: 1.72 requests/s, 2202.42 total tokens/s, 440.48 output tokens/s
Total num prompt tokens:  1024000
Total num output tokens:  256000

PR: calculate_kv_scales=False | enforce_eager=False


Throughput: 1.73 requests/s, 2211.37 total tokens/s, 442.27 output tokens/s
Total num prompt tokens:  1024000
Total num output tokens:  256000

PR: calculate_kv_scales=True | enforce_eager=False


Throughput: 1.73 requests/s, 2210.78 total tokens/s, 442.16 output tokens/s
Total num prompt tokens:  1024000
Total num output tokens:  256000

PR: calulate_kv_scales=True | enforce_eager=True

Throughput: 1.72 requests/s, 2200.57 total tokens/s, 440.11 output tokens/s
Total num prompt tokens:  1024000
Total num output tokens:  256000

E2E GSM8K Accuracy Results

Main: calulate_kv_scales=False | enforce_eager=False


|Tasks|Version|     Filter     |n-shot|  Metric   |   |Value|   |Stderr|
|-----|------:|----------------|-----:|-----------|---|----:|---|-----:|
|gsm8k|      3|flexible-extract|     5|exact_match|↑  | 0.84|±  |0.0368|
|     |       |strict-match    |     5|exact_match|↑  | 0.83|±  |0.0378|```

Main: calulate_kv_scales=True | enforce_eager=False

|Tasks|Version|     Filter     |n-shot|  Metric   |   |Value|   |Stderr|
|-----|------:|----------------|-----:|-----------|---|----:|---|-----:|
|gsm8k|      3|flexible-extract|     5|exact_match|↑  | 0.85|±  |0.0359|
|     |       |strict-match    |     5|exact_match|↑  | 0.86|±  |0.0349|

Main: calulate_kv_scales=True | enforce_eager=True

|Tasks|Version|     Filter     |n-shot|  Metric   |   |Value|   |Stderr|
|-----|------:|----------------|-----:|-----------|---|----:|---|-----:|
|gsm8k|      3|flexible-extract|     5|exact_match|↑  | 0.88|±  |0.0327|
|     |       |strict-match    |     5|exact_match|↑  | 0.89|±  |0.0314|

PR: calculate_kv_scales=False | enforce_eager=False


|Tasks|Version|     Filter     |n-shot|  Metric   |   |Value|   |Stderr|
|-----|------:|----------------|-----:|-----------|---|----:|---|-----:|
|gsm8k|      3|flexible-extract|     5|exact_match|↑  | 0.86|±  |0.0349|
|     |       |strict-match    |     5|exact_match|↑  | 0.86|±  |0.0349|

PR: calculate_kv_scales=True | enforce_eager=False


|Tasks|Version|     Filter     |n-shot|  Metric   |   |Value|   |Stderr|
|-----|------:|----------------|-----:|-----------|---|----:|---|-----:|
|gsm8k|      3|flexible-extract|     5|exact_match|↑  | 0.86|±  |0.0349|
|     |       |strict-match    |     5|exact_match|↑  | 0.86|±  |0.0349|

PR: calulate_kv_scales=True | enforce_eager=True

|Tasks|Version|     Filter     |n-shot|  Metric   |   |Value|   |Stderr|
|-----|------:|----------------|-----:|-----------|---|----:|---|-----:|
|gsm8k|      3|flexible-extract|     5|exact_match|↑  | 0.89|±  |0.0314|
|     |       |strict-match    |     5|exact_match|↑  | 0.90|±  |0.0302|

Signed-off-by: adabeyta <aabeyta@redhat.com>
@mergify mergify Bot added the v1 label Oct 21, 2025

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request correctly addresses a bug where calculated KV scales were not being applied during attention. The fix introduces a mechanism to calculate these scales on the first forward pass and then disables subsequent calculations by adding a state flag kv_scales_calculated to GPUModelRunner. The logic is sound and the changes are well-targeted. I've included one suggestion to refactor the implementation slightly, which will improve maintainability by removing a small piece of redundant logic.

Comment thread vllm/v1/worker/gpu_model_runner.py Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread vllm/v1/worker/gpu_model_runner.py Outdated

@ProExpertProg ProExpertProg left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thx for the fix, I got a suggestion to further simplify the logic

Comment thread vllm/v1/attention/backends/utils.py Outdated
Comment thread vllm/v1/worker/gpu_model_runner.py Outdated
Comment thread vllm/v1/worker/gpu_model_runner.py Outdated
Comment thread vllm/v1/worker/gpu_model_runner.py Outdated
Comment thread vllm/v1/worker/gpu_model_runner.py Outdated
@ProExpertProg

Copy link
Copy Markdown
Collaborator

Could you just sanity check with a deepseek model that it works with MLA too?

@ProExpertProg ProExpertProg left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Seems like there's a few more paths left over

Comment thread vllm/v1/worker/gpu_model_runner.py Outdated
Comment thread vllm/attention/layer.py Outdated
Comment thread vllm/attention/layer.py Outdated
Signed-off-by: adabeyta <aabeyta@redhat.com>
Signed-off-by: adabeyta <aabeyta@redhat.com>
Comment thread vllm/attention/layer.py Outdated
@ProExpertProg ProExpertProg added the ready ONLY add when PR is ready to merge/full CI is needed label Nov 3, 2025
@mergify mergify Bot added the ci/build label Nov 4, 2025
Comment thread vllm/attention/layer.py Outdated
…ntion

Signed-off-by: adabeyta <aabeyta@redhat.com>
@mergify

mergify Bot commented Nov 7, 2025

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @adabeyta.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Nov 7, 2025
Signed-off-by: adabeyta <aabeyta@redhat.com>

# Conflicts:
#	.buildkite/test-pipeline.yaml
@mergify mergify Bot removed the needs-rebase label Nov 7, 2025

@ProExpertProg ProExpertProg left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just a CI file merge note

Comment thread .buildkite/test-pipeline.yaml Outdated
Comment thread vllm/attention/layer.py
Signed-off-by: adabeyta <aabeyta@redhat.com>
@ProExpertProg
ProExpertProg enabled auto-merge (squash) November 10, 2025 16:34
@ProExpertProg
ProExpertProg enabled auto-merge (squash) November 10, 2025 23:38
@ProExpertProg
ProExpertProg merged commit a5a790e into vllm-project:main Nov 10, 2025
49 checks passed
@ywang96 ywang96 added this to the v0.11.1 milestone Nov 13, 2025
khluu pushed a commit that referenced this pull request Nov 16, 2025
khluu pushed a commit that referenced this pull request Nov 16, 2025
Signed-off-by: adabeyta <aabeyta@redhat.com>
(cherry picked from commit a5a790e)
devpatelio pushed a commit to SumanthRH/vllm that referenced this pull request Nov 29, 2025
mystous pushed a commit to mystous/vllm_hybrid that referenced this pull request May 10, 2026
my-other-github-account pushed a commit to my-other-github-account/vllm that referenced this pull request May 15, 2026
my-other-github-account pushed a commit to my-other-github-account/vllm that referenced this pull request May 15, 2026
…project#27232)

Signed-off-by: adabeyta <aabeyta@redhat.com>
(cherry picked from commit a5a790e)
0826joyce pushed a commit to 0826joyce/vllm-serving-optimization that referenced this pull request May 19, 2026
plasticchris pushed a commit to plasticchris/vllm that referenced this pull request Jul 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci/build ready ONLY add when PR is ready to merge/full CI is needed v1

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants