⚠️ Please check that this feature request hasn't been suggested before.
🔖 Feature description
The default huggingface moe implementation is inefficient due to for loop on experts, which causes low utilization when training models like Qwen3 30B-A3B, we might want a drop in kernel patch for more efficient moe calculation
A previous issue on this #930
✔️ Solution
Integrating solutions like https://github.com/databricks/megablocks/
❓ Alternatives
No response
📝 Additional Context
No response
Acknowledgements
🔖 Feature description
The default huggingface moe implementation is inefficient due to for loop on experts, which causes low utilization when training models like Qwen3 30B-A3B, we might want a drop in kernel patch for more efficient moe calculation
A previous issue on this #930
✔️ Solution
Integrating solutions like https://github.com/databricks/megablocks/
❓ Alternatives
No response
📝 Additional Context
No response
Acknowledgements