v0.23.0.post1 - 2026.09.21
This is the first post release of vLLM Ascend v0.23.0. It includes the fixes, dependency updates, CI changes, and documentation updates merged into the v0.23.0 release branch after the v0.23.0 tag. Please follow the official documentation to get started.
Bug Fixes
- Fixed stale KV-cache writes during DP-aligned dummy runs by invalidating per-KV-group slot mappings before attention metadata construction. #15362
- Fixed MTP overlay prefix-cache precision on Atlas 300I DUO and kept W8A8 MXFP8 transformed buffers stable across RL weight reloads in ACL Graph mode. #14336 #13905
Other Changes
- Pinned the KV Pool dependencies to
memfabric_hybrid==1.2.0andmemcache_hybrid==1.2.0. #14352 - Consolidated installation guidance, GLM-5/5.2 and Kimi-K3 deployment instructions, PD and 310P notes, release metadata, navigation titles, English comments, and Chinese translations. #15449 #14242 #14338 #14634 #14698 #14713 #14906 #15998 #16106 #14382 #14387 #14436 #14579 #14684 #15534 #16202
- Added release-branch nightly and weekly model configurations and installed
concurrent-log-handlerin release images. #14559 #14645 #14739 #16191