Divide and conquer: Scalable performance and energy in MCM GPUs
By Mario Ibáñez Bolado, Borja Pérez Pavón, Jose Luis Bosque Orero, Julio Ramón Beivide
Universidad de Cantabria, Spain

Abstract
Multi-chip-module (MCM) GPUs offer a promising path to scale compute capability beyond monolithic designs by integrating multiple chiplets on a common package. However, the impact of disaggregation on performance scalability and energy consumption remains underexplored. The design space grows rapidly across dimensions such as SMs per chiplet, chiplet count, and interconnection network. The inter-chiplet network is particularly critical, as it determines whether additional compute resources translate into performance gains. This limited understanding leaves industry and research without clear guidance on the performance and energy trade-offs of MCM GPU scaling. In this work, we investigate whether distributing compute and memory capability across multiple chiplets offers a more scalable alternative to concentrating resources. We quantify their effects on performance, energy, and efficiency and examine how inter-chiplet topology influences scalability at different system sizes. Our results demonstrate that a 16 chiplet Torus configuration with 256 SMs delivers a remarkable 2.40× performance improvement over a state-of-the-art MCM architecture with the same compute capability, while simultaneously reducing energy consumption by 4.45×. These substantial gains provide evidence that disaggregation is a first-order architectural factor and will be critical to unlocking the performance and energy-efficiency potential of next-generation GPUs.
To read the full article, click here
Related Chiplet
- FlexGen Multi-Die Smart Network-on-Chip (NoC) IP
- Ncore Multi-Die Interconnect IP
- Integrated voltage regulator (IVR) chiplet
- High-performance connectivity chiplets
- eFPGA Chiplet
Related Technical Papers
- Signal Integrity Challenges in Chiplet-Based Designs: Addressing Performance and Security
- A Time-Encoded Analog Photonic Interposer for Energy-Efficient Integration of Analog Vision Sensors and Analog Accelerators
- The Next Frontier in Semiconductor Innovation: Chiplets and the Rise of 3D-ICs
- Codesign of quantum error-correcting codes and modular chiplets in the presence of defects
Latest Technical Papers
- Divide and conquer: Scalable performance and energy in MCM GPUs
- A Composable AI-Accelerated Iterative Solver for 3D-IC Thermal Modeling
- QBX: A Compiler for 2-local Qubit Hamiltonian Simulation on Quantum Chiplets
- U-Can-Inject-Errors (UCIe): A Protocol-Aware Hardware Trojan for FPGA Chiplet Links
- The Power of Indirection: Scaling Switches Beyond Silicon Boundaries