Skip to content

Integration of fused moe kernel (e.g., megablocks) for efficient moe training #3155

Description

@zinccat

⚠️ Please check that this feature request hasn't been suggested before.

  • I searched previous Ideas in Discussions didn't find any similar feature requests.
  • I searched previous Issues didn't find any similar feature requests.

🔖 Feature description

The default huggingface moe implementation is inefficient due to for loop on experts, which causes low utilization when training models like Qwen3 30B-A3B, we might want a drop in kernel patch for more efficient moe calculation
A previous issue on this #930

✔️ Solution

Integrating solutions like https://github.com/databricks/megablocks/

❓ Alternatives

No response

📝 Additional Context

No response

Acknowledgements

  • My issue title is concise, descriptive, and in title casing.
  • I have searched the existing issues to make sure this feature has not been requested yet.
  • I have provided enough information for the maintainers to understand and evaluate this request.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions