Preprint / Version 1

Cooling-Aware Routing in Mixture-of-Experts Models: Optimizing Expert Assignment Under Routing-Deviation Constraints

##article.authors##

  • Shreyasi Saxena -

DOI:

https://doi.org/10.58445/rars.4177

Keywords:

Mixture-of-Experts, MoE Routing, Switch Transformer, expert routing, load balancing, GPU workload, direct-to-chip liquid-cooling, thermal management, cooling-aware, optimization, router deviation, expert placement, pump power

Abstract

Mixture-of-Experts (MoE) models reduce computational demand by only activating a subset of available experts for each input token, but uneven expert selection often produces imbalanced workloads when experts are distributed across physical accelerators. This study investigates whether token-level routing flexibility can be used to reduce modeled direct-to-chip liquid-cooling demand while limiting departure from the model’s original router preferences. Routing data were collected from the Switch-Base-8 using five separately seeded 1,000-document samples from the C4 validation dataset, including 978,748 non-padding token positions and 5,872,488 token-layer routing decisions. Expert assignments were mapped only to an 8-GPU configuration and connected to a reduced-order thermal and hydraulic model. Next, a constrained optimization was used to redistribute selected token assignments while minimizing peak GPU workload and routing deviation. Across the five samples, baseline maximum GPU workload was approximately 21% above the equal-share reference. At a target corresponding to 75% of the maximum achievable workload balancing, an average on just 3.16% of token-layer assignments were rerouted, with the mean routing deviation coming to 0.009857, while the maximum GPU workload decreased by 12.96%. The result was reproducible across separately seeded samples and stayed present under alternative expert-to-GPU placements. Under modeled hydraulic assumptions, the corresponding mean pump-power reduction ranged from 12.96% to 24.18%. These results indicate that token-level MoE routing flexibility may provide a useful mechanism for reducing modeled peak routing-dependent cooling demand, even though hardware validation and direct evaluation of model-quantity effects remain necessary.

References

Stojkovic J, Zhang C, Goiri Í, Choukse E, Qiu H, Fonseca R, et al. TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms. Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, pp. 1266-1281, 2025. https://doi.org/10.1145/3676641.3716025

Heydari A, Soud A, Tradat M, Soud Q, Eslami B, Shahi P, et al. Experimental Evaluation of Cooling Loop Configurations for Server-Level Direct-to-Chip Liquid Cooling Systems. J Electron Packag, 1-25, 2026. Paper No. EP-26-1034. https://doi.org/10.1115/1.4072567

Fedus W, Zoph B, Shazeer N. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. J Mach Learn Res, 23(120): 1-39, 2022. https://doi.org/10.48550/arXiv.2101.03961

Liu J, Tang P, Wang W, Ren Y, Hou X, Heng PA, et al. A Survey on Inference Optimization Techniques for Mixture of Experts Models. ACM Comput Surv, 58(10): Article 247, 1-37, 2026. https://doi.org/10.1145/3794845

Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. J Mach Learn Res, 21(140): 1-67, 2020. https://doi.org/10.48550/arXiv.1910.10683

Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, et al. Transformers: State-of-the-Art Natural Language Processing. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38-45, 2020. https://doi.org/10.18653/v1/2020.emnlp-demos.6

Heydari A, Gharaibeh AR, Tradat M, Soud Q, Manaserh Y, Radmard V, et al. Experimental evaluation of direct-to-chip cold plate liquid cooling for high-heat-density data centers. Appl Therm Eng, 239: 122122, 2024. https://doi.org/10.1016/j.applthermaleng.2023.122122

Rennels D. Pipe Flow: A Practical and Comprehensive Guide. 2nd ed., John Wiley & Sons, Hoboken, NJ, USA, 2022. https://doi.org/10.1002/9781119756460

Hwang C, Cui W, Xiong Y, Yang Z, Liu Z, Hu H, et al. Tutel: Adaptive Mixture-of-Experts at Scale. Proc Mach Learn Syst, 5: 269-287, 2023. https://doi.org/10.48550/arXiv.2206.03382

Go S, Mahajan D. MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing. arXiv preprint arXiv:2502.06643, 2025. https://doi.org/10.48550/arXiv.2502.06643

Chen Q, Chen X, Huang K. SiftMoE: Similarity-Aware Energy-Efficient Expert Selection for Wireless Distributed MoE Inference. arXiv preprint arXiv:2603.23888, 2026. https://doi.org/10.48550/arXiv.2603.23888

Downloads

Posted

2026-09-20