🥝flash-linear-attention + NPU: fla-org/flash-linear-attention#942
🍇linkedin/Liger-Kernel + NPU: linkedin/Liger-Kernel#969
🍈Ascend/docs: https://ascend.github.io/docs
🍉verl-project/verl + NPU: verl-project/verl#900
🥝flash-linear-attention + NPU: fla-org/flash-linear-attention#942
🍇linkedin/Liger-Kernel + NPU: linkedin/Liger-Kernel#969
🍈Ascend/docs: https://ascend.github.io/docs
🍉verl-project/verl + NPU: verl-project/verl#900
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Efficient Triton Kernels for LLM Training
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Forked from fla-org/flash-linear-attention
🚀 Efficient implementations for emerging model architectures
Python 1
🚀 Efficient implementations for emerging model architectures