Efficient MoE inference

Dynamic gating, expert buffering, and load balancing for mixture-of-experts models.

Mixture-of-experts models increase model capacity while keeping computation sparse, but their memory requirements and communication patterns make deployment challenging. This project studies those bottlenecks and develops dynamic gating, expert buffering, and expert load balancing to improve inference efficiency.

The work appeared at NeurIPS 2024.

Read the paper · Explore the code