ML Engineer · Google

Haiyang Huang

Research on efficient
machine learning systems.

Efficient ML systems · Scalable inference · Interpretable learning

Haiyang Huang
Efficient systems. Interpretable learning.

About me

I’m an ML Engineer at Google. I graduated with a PhD in Computer Science from Duke University, where I was advised by Cynthia Rudin and Benjamin C. Lee.

During my PhD, I studied efficient machine learning systems and scalable inference, including memory use and inference efficiency for mixture-of-experts models. My doctoral research also covered interpretable machine learning and dimensionality reduction for data visualization.

Before Duke, I earned a bachelor’s degree in Mathematics and Computer Science at the University of Michigan, where I worked with Jenna Wiens.

01

Efficient ML systems

Improving large-scale inference through dynamic gating, expert buffering, and load balancing for mixture-of-experts models.

Explore the work
02

Interpretable learning

Understanding visual concepts and building models whose decisions we can inspect.

03

Data visualization

Revealing local and global structure in high-dimensional data.

Selected research

All publications
  1. AISTATS
    The Rashomon Effect for Visualizing High-Dimensional Data
    Yiyang Sun , Haiyang Huang, Gaurav Rajesh Parikh , and 1 more author
    In Proceedings of the 29th International Conference on Artificial Intelligence and Statistics , 2026
  2. AAAI
    Dimension Reduction with Locally Adjusted Graphs
    Yingfan Wang , Yiyang Sun , Haiyang Huang, and 1 more author
    In Proceedings of the AAAI Conference on Artificial Intelligence , 2025
  3. NeurIPS
    Toward Efficient Inference for Mixture of Experts
    Haiyang Huang, Newsha Ardalani , Anna Sun , and 5 more authors
    In Advances in Neural Information Processing Systems , 2024
  4. NeurIPS
    Navigating the Effect of Parametrization for Dimensionality Reduction
    Haiyang Huang, Yingfan Wang , and Cynthia Rudin
    In Advances in Neural Information Processing Systems , 2024
  5. JMLR
    Understanding How Dimension Reduction Tools Work: An Empirical Approach to Deciphering t-SNE, UMAP, TriMap, and PaCMAP for Data Visualization
    Yingfan Wang , Haiyang Huang, Cynthia Rudin , and 1 more author
    Journal of Machine Learning Research, 2021

Updates

Archive
Apr 01, 2026 The Rashomon Effect for Visualizing High-Dimensional Data is accepted at AISTATS 2026.
Apr 11, 2025 Our LocalMAP paper is published in the AAAI 2025 proceedings.
Nov 10, 2024 Two of my papers were accepted at NeurIPS 2024: efficient MoE inference and parametric dimensionality reduction.
May 06, 2022 I am thrilled to join Facebook AI Research (FAIR) as a research intern starting at the middle of May!

Get in touch

Let’s connect.

Reach out on LinkedIn for research conversations and collaborations.