Characterizing and Mitigating Performance Variability in Accelerator-Rich Systems

Talk
Matt Sinclair
University of Wisconsin-Madison
Time: 
10.21.2026 15:30 to 16:30
Location: 

IRB 4105

In recent years, to reach performance goals modern computing systems are increasingly turning to using large numbers of compute accelerators, which offer greater power efficiency and thus enable higher performance within a constrained power budget. However, using accelerators increases heterogeneity at multiple levels, including the architecture, resource allocation, competing user needs, and manufacturing variability. Accordingly, current and future systems need to efficiently handle many simultaneous jobs while balancing PM and multiple levels of heterogeneity. In recent work, we have demonstrated the extent of this variability in modern accelerator-rich systems (SC'22) and shown how to embrace variability in cluster-level job schedulers (SC'24). This work significantly improves the efficiency of modern systems for a range of ML workloads. However, scheduling jobs more efficiently at the software and runtime layers is limited in its ability to quickly, dynamically change policies as cluster conditions evolve. A major limiter to further improving efficiency is the lack of standards for exposing power information in modern accelerators (SIGMETRICS'26). Thus, for future systems we propose to build on the insights generated by our optimizations for current systems, and apply co-design that makes the hardware, software, and runtime layers aware of the variance in the systems.