Skip to content

FEAT: support bidirectional Xavier P/D between NVIDIA and MLX - #5641

Open
qinxuye wants to merge 3 commits into
xorbitsai:mainfrom
qinxuye:feat/xavier-gpu-mlx-pd
Open

qinxuye wants to merge 3 commits into
xorbitsai:mainfrom
qinxuye:feat/xavier-gpu-mlx-pd

Conversation

@qinxuye

@qinxuye qinxuye commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Adds bidirectional NVIDIA/Metal P/D: vLLM → MLX, SGLang → MLX, MLX → vLLM and MLX → SGLang, including streaming/non-streaming chat and completions. CUDA/Metal handoff uses canonical FP16 64-token pages through bounded CPU staging and actor RPC. MLX prefill owns pages on its worker under a 512 MiB in-flight budget; SGLang decode commits full prompt KV and P’s first token, while vLLM/MLX decode recomputes the last prompt token. Cancellation drains CUDA operations before releasing engine slots, and termination/failed launch removes worker-local source actors.

Dependencies

Based on main at 4969faffd after #5631, #5638 and #5639 merged; no unmerged PR dependencies. The PR now contains only the NVIDIA/MLX host-handoff extension and its integration tests/docs.

Initial scope: identical original unquantized Qwen2/Qwen3/Llama full-attention checkpoints and tokenizer assets, effective FP16 weights, TP=PP=DP=1 and n=1, without LoRA or multimodal/hybrid caches. Per-replica engine_config accepts worker-local model paths/formats. Weight/tokenizer/context/token-ID mismatches and transfer failures raise errors. Supervisor and worker/model environments must use matching Python and RPC serialization versions. Documentation and all nine translations are updated; the benchmark runs either or both directions on an existing cluster.

Final performance

Median TTFT (ms):

Prompt tokens MLX local vLLM local SGLang local vLLM P → MLX D SGLang P → MLX D MLX P → vLLM D MLX P → SGLang D
51 50.5 75.3 62.5 188.8 235.7 365.9 296.5
184 53.0 76.3 62.4 211.2 254.6 382.2 321.6
744 74.3 82.2 71.1 283.2 345.1 519.3 447.9
2844 170.4 116.0 97.7 587.0 609.1 1013.7 898.6

Qwen2.5-0.5B-Instruct, FP16, RTX 3090 Ti and M5 Pro/64 GB, vLLM 0.28.0 / SGLang 0.5.21 / mlx-lm 0.31.3. Context 8,192, eager CUDA execution, C1, greedy streaming, 32 output tokens, five samples per prompt length after warmup, distinct early prefixes and no complete-page cache reuse. Standalone vLLM uses its default 16-token blocks; P/D uses canonical 64-token pages. Median end-to-end TTFT includes routing, prefill, CPU/network copies and decode. Measured LAN payload throughput is approximately 85 MB/s in both directions. Each retained GPU sample passed the external ownership monitor; runs with observed measurement interference are retried. These P/D combinations have no TTFT advantage over standalone decode on this small model and current link. The table retains the previous final measurements; this rebase reruns functional/lifecycle validation and does not claim a new full performance comparison.

Validation

  • Latest-head Python/Metal regression sweep: 1,159 passed, 49 skipped; whole-repository pre-commit passed.

  • Latest-head real LAN chat/completions and streaming/non-streaming requests passed in all four directions, including 64/65/66/128/129/130/193-token boundaries. Completed imports are verified by imported_tokens / host_bytes, with zero gpu_bytes and no active handoffs after completion. The merged vLLM/SGLang cancellation retry and page-allocation fixes remain in place; host imports also cover minimum/full target allocation. Invalid per-replica weight formats fail before directory/source/model creation.

  • Lifecycle: parallel requests, stream disconnect, cancellation during CPU handoff, empty source/transfer state and successful fresh requests after cancellation, in both directions.

  • In the previous full performance run, same-prompt greedy text matched standalone FP16 MLX in 18/20 vLLM → MLX, 19/20 SGLang → MLX, 17/20 MLX → vLLM and 19/20 MLX → SGLang final samples. Some FP16 continuations differ between engines; cross-engine greedy text equality is not guaranteed.

  • All nine changed PO catalogs validated and compiled; P/D and Xavier pages rendered in English and every maintained language (20 pages), with translated text verified.

  • CI from the previous head: Windows shutdown-timing failures are already fixed on main; Linux 3.13 encountered Hugging Face HTTP 429 during model downloads. Fresh CI will run on the rebased head.

@qinxuye
qinxuye force-pushed the feat/xavier-gpu-mlx-pd branch from d3294d9 to 52cc6b2 Compare October 7, 2026 03:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants