Repository navigation
Conversation
qinxuye
force-pushed
the
feat/xavier-gpu-mlx-pd
branch
from
October 7, 2026 03:56
d3294d9 to
52cc6b2
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds bidirectional NVIDIA/Metal P/D: vLLM → MLX, SGLang → MLX, MLX → vLLM and MLX → SGLang, including streaming/non-streaming chat and completions. CUDA/Metal handoff uses canonical FP16 64-token pages through bounded CPU staging and actor RPC. MLX prefill owns pages on its worker under a 512 MiB in-flight budget; SGLang decode commits full prompt KV and P’s first token, while vLLM/MLX decode recomputes the last prompt token. Cancellation drains CUDA operations before releasing engine slots, and termination/failed launch removes worker-local source actors.
Dependencies
Based on
mainat4969faffdafter #5631, #5638 and #5639 merged; no unmerged PR dependencies. The PR now contains only the NVIDIA/MLX host-handoff extension and its integration tests/docs.Initial scope: identical original unquantized Qwen2/Qwen3/Llama full-attention checkpoints and tokenizer assets, effective FP16 weights, TP=PP=DP=1 and
n=1, without LoRA or multimodal/hybrid caches. Per-replicaengine_configaccepts worker-local model paths/formats. Weight/tokenizer/context/token-ID mismatches and transfer failures raise errors. Supervisor and worker/model environments must use matching Python and RPC serialization versions. Documentation and all nine translations are updated; the benchmark runs either or both directions on an existing cluster.Final performance
Median TTFT (ms):
Qwen2.5-0.5B-Instruct, FP16, RTX 3090 Ti and M5 Pro/64 GB, vLLM 0.28.0 / SGLang 0.5.21 / mlx-lm 0.31.3. Context 8,192, eager CUDA execution, C1, greedy streaming, 32 output tokens, five samples per prompt length after warmup, distinct early prefixes and no complete-page cache reuse. Standalone vLLM uses its default 16-token blocks; P/D uses canonical 64-token pages. Median end-to-end TTFT includes routing, prefill, CPU/network copies and decode. Measured LAN payload throughput is approximately 85 MB/s in both directions. Each retained GPU sample passed the external ownership monitor; runs with observed measurement interference are retried. These P/D combinations have no TTFT advantage over standalone decode on this small model and current link. The table retains the previous final measurements; this rebase reruns functional/lifecycle validation and does not claim a new full performance comparison.
Validation
Latest-head Python/Metal regression sweep: 1,159 passed, 49 skipped; whole-repository pre-commit passed.
Latest-head real LAN chat/completions and streaming/non-streaming requests passed in all four directions, including 64/65/66/128/129/130/193-token boundaries. Completed imports are verified by
imported_tokens/host_bytes, with zerogpu_bytesand no active handoffs after completion. The merged vLLM/SGLang cancellation retry and page-allocation fixes remain in place; host imports also cover minimum/full target allocation. Invalid per-replica weight formats fail before directory/source/model creation.Lifecycle: parallel requests, stream disconnect, cancellation during CPU handoff, empty source/transfer state and successful fresh requests after cancellation, in both directions.
In the previous full performance run, same-prompt greedy text matched standalone FP16 MLX in 18/20 vLLM → MLX, 19/20 SGLang → MLX, 17/20 MLX → vLLM and 19/20 MLX → SGLang final samples. Some FP16 continuations differ between engines; cross-engine greedy text equality is not guaranteed.
All nine changed PO catalogs validated and compiled; P/D and Xavier pages rendered in English and every maintained language (20 pages), with translated text verified.
CI from the previous head: Windows shutdown-timing failures are already fixed on
main; Linux 3.13 encountered Hugging Face HTTP 429 during model downloads. Fresh CI will run on the rebased head.