Skip to content

vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled request id — cross-request integrity break and induced errors on `/score` and `/rerank`

Moderate severity GitHub Reviewed Published Sep 23, 2026 in vllm-project/vllm • Updated Oct 5, 2026

Package

pip vllm (pip)

Affected versions

< 0.30.0

Patched versions

0.30.0

Description

Affected

  • Ecosystem / package: pip / vllm
  • Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the flash late-interaction query cache reaches.

Summary

On late-interaction /score and /rerank deployments with flash late interaction enabled (the default for supported models), the worker caches per-request query embeddings under a key derived from the caller-controlled X-Request-Id header. A second concurrent request that reuses the victim's header value replaces the victim's cached query embedding before document scoring — so the victim's documents are scored against the attacker's query. Because the data-parallel router pins all requests sharing a cache key to the same engine, the collision is deterministic for an attacker who reuses the victim's X-Request-Id. Depending on timing, one request can also consume the shared use counter and force the other request into a late-interaction cache-miss error.

This is a remotely reachable, request-controlled cross-request integrity break on the standard scoring and reranking endpoints. It requires only that flash late interaction be enabled, which is the default for supported models.

Affected code

Links pinned to the confirmed commit 752a3a504485 (v0.25.1):

The caller-controlled header enters as the request id, and the flash late-interaction path derives the worker cache key directly from it:

# vllm/entrypoints/serve/engine/serving.py Lines 116-126
    @staticmethod
    def _base_request_id(
        raw_request: Request | None, default: str | None = None
    ) -> str | None:
        """Pulls the request id to use from a header, if provided"""
        if raw_request is not None and (
            (req_id := raw_request.headers.get("X-Request-Id")) is not None
        ):
            return req_id

        return random_uuid() if default is None else default
# vllm/entrypoints/pooling/scoring/serving.py Lines 207-212
        n_queries = ctx.n_queries
        n_docs = len(ctx.engine_inputs) - n_queries
        query_engine_inputs = ctx.engine_inputs[:n_queries]

        query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
        query_uses = [n_docs if n_queries == 1 else 1] * n_queries

The worker then stores and reads the query embedding under that string with no check that the reader owns the entry — a colliding key returns another request's cached query, or (once the use counter is exhausted) raises a cache-miss error:

# vllm/v1/worker/gpu/pool/late_interaction_runner.py Lines 91-107
            if mode == LATE_INTERACTION_MODE_CACHE_QUERY:
                assert query_uses is not None
                # `output` can be a view into the current step's hidden-states
                # buffer, so clone it before storing across scheduling steps.
                self._query_cache[query_key] = output.clone()
                self._query_uses[query_key] = query_uses
                outputs[i] = torch.zeros((), device=output.device, dtype=torch.float32)
                continue

            if mode == LATE_INTERACTION_MODE_SCORE_DOC:
                query_output = self._query_cache.get(query_key)
                if query_output is None:
                    raise ValueError(
                        "late-interaction query cache miss for key "
                        f"{query_key!r}. Ensure query requests are executed "
                        "before their paired document requests."
                    )

The bug is specific to the flash late-interaction path. Non-flash late-interaction scoring computes MaxSim directly from one request's in-memory outputs and does not create a cross-request worker cache key.

Impact

A network client of the standard scoring API can, on a flash late-interaction /score or /rerank deployment:

  1. Corrupt another user's results — by reusing the victim's X-Request-Id, the attacker's query embedding overwrites the victim's cached entry, so the victim's documents are scored against the attacker's query (a cross-request integrity break).
  2. Induce errors — depending on timing, one request consumes the shared use counter and forces the other request into a late-interaction cache-miss error.

Both consequences follow deterministically from reusing the victim's header value, because same-key work is pinned to one engine. This is reachable through normal request handling and does not depend on any trusted inter-node network.

Suggested Fix

Derive the flash late-interaction query-cache key from a server-generated, unforgeable per-request identifier (a random_uuid() namespace) rather than the caller-supplied X-Request-Id, and thread that key through the PoolingServeContext to the doc-scoring pass so both passes reuse the same key and a caller cannot address another request's cache entry.

In vllm/entrypoints/pooling/scoring/serving.py, the encode-queries pass mints a fresh namespace and stashes the keys on the context:

-        query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
+        query_namespace = random_uuid()
+        query_keys = [
+            f"late-interaction-{query_namespace}-query-{i}" for i in range(n_queries)
+        ]
+        ctx.late_interaction_query_keys = query_keys

and the encode-docs pass reads those stored keys instead of re-deriving them from ctx.request_id:

-        query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
+        query_keys = ctx.late_interaction_query_keys
+        if query_keys is None:
+            raise RuntimeError("Late-interaction query keys were not initialized.")

This requires adding the late_interaction_query_keys: list[str] | None = None field to PoolingServeContext (vllm/entrypoints/pooling/typing.py). Because the namespace is a server-generated UUID, colliding X-Request-Id values no longer produce a shared cache key; a regression test asserting exactly that (colliding request ids yield distinct query-cache keys) accompanies the change.

Credit

Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)

This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.


Proposed fix: a fix for this issue is proposed in a public pull request: vllm-project/vllm#51445

References

@jperezdealgaba jperezdealgaba published to vllm-project/vllm Sep 23, 2026
Published to the GitHub Advisory Database Oct 5, 2026
Reviewed Oct 5, 2026
Last updated Oct 5, 2026

Severity

Moderate

CVSS overall score

This score calculates overall vulnerability severity from 0 to 10 and is based on the Common Vulnerability Scoring System (CVSS).
/ 10

CVSS v3 base metrics

Attack vector
Network
Attack complexity
High
Privileges required
Low
User interaction
None
Scope
Unchanged
Confidentiality
None
Integrity
Low
Availability
Low

CVSS v3 base metrics

Attack vector: More severe the more the remote (logically and physically) an attacker can be in order to exploit the vulnerability.
Attack complexity: More severe for the least complex attacks.
Privileges required: More severe if no privileges are required.
User interaction: More severe when no user interaction is required.
Scope: More severe when a scope change occurs, e.g. one vulnerable component impacts resources in components beyond its security scope.
Confidentiality: More severe when loss of data confidentiality is highest, measuring the level of data access available to an unauthorized user.
Integrity: More severe when loss of data integrity is the highest, measuring the consequence of data modification possible by an unauthorized user.
Availability: More severe when the loss of impacted component availability is highest.
CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:N/I:L/A:L

EPSS score

Exploit Prediction Scoring System (EPSS)

This score estimates the probability of this vulnerability being exploited within the next 30 days. Data provided by FIRST.
(9th percentile)

Weaknesses

Authorization Bypass Through User-Controlled Key

The system's authorization functionality does not prevent one user from gaining access to another user's data or record by modifying the key value identifying the data. Learn more on MITRE.

CVE ID

CVE-2026-105755

GHSA ID

GHSA-2phq-3phc-84px

Source code

Credits

Loading Checking history
See something to contribute? Suggest improvements for this vulnerability.