Fengshui: Demystifying Chiplet Ecosystem and Bespoke Neural Network Accelerator Codesign
By Haoran Jin, Jirong Yang, Zhiheng Zhang, Justin Shin, Barry Lyu, Kangqi Zhang, Yunpeng Liu, Nathan Bleier
University of Michigan, Ann Arbor, MI, USA

Abstract
Modern ML workloads, with stringent latency and energy constraints, are increasingly hard to run efficiently on homogeneous commodity hardware. We argue that operatorlevel disaggregation—tailoring microarchitecture, batching, and memory hierarchy to each operator—is essential to overcome these limitations, though the resulting highly bespoke accelerators incur prohibitive Non-Recurring Engineering (NRE) costs. Chiplet-based integration amortizes NRE across applications, but choosing which chiplets to build and how to compose them into accelerators is circularly dependent—a chiplet pool’s value depends on the constructed accelerators, while accelerator quality is constrained by available chiplets.
This paper introduces Fengshui, a chiplet ecosystem and accelerator co-design framework that jointly optimizes chiplet pool composition and bespoke application-specific integrated circuit (BASIC) design. Fengshui constructs BASICs through operatorlevel disaggregation, co-exploring chiplet and memory heterogeneity, tensor fusion, and pipeline/tensor/expert parallelism with place-and-route validation for physical implementability.
With just 8 strategically selected chiplets, encompassing network switches, processing-in-memory units, and accelerators with diverse microarchitectures, Fengshui-generated BASICs achieve 48.5%, 88.1%, 93.0%, and 97.8% reductions in energy, energycost product (energy×$), energy-delay product (EDP), and energy-delay-cost product (EDP×$) over homogeneous accelerators, while scoring within 4.1% of unconstrained heterogeneous designs across diverse neural networks. For datacenter MoE and dense LLM serving, Fengshui reduces prefill energy and energy×$ by up to 16.8% and 28.7%, respectively; for edge autonomous vehicle perception, it achieves 12.0% energy and 23.6% energy×$ reductions under real-time latency constraints.
To read the full article, click here
Related Chiplet
- Integrated voltage regulator (IVR) chiplet
- High-performance connectivity chiplets
- eFPGA Chiplet
- DPIQ Tx PICs
- IMDD Tx PICs
Related Technical Papers
- DISTIL: A Distributed Spiking Neural Network Accelerator on 2.5D Chiplet Systems
- Understanding and Profiling the Accelerator Chiplet Network Using PingPoint
- Mozart: A Chiplet Ecosystem-Accelerator Codesign Framework for Composable Bespoke Application Specific Integrated Circuits
- Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet Accelerators
Latest Technical Papers
- Fengshui: Demystifying Chiplet Ecosystem and Bespoke Neural Network Accelerator Codesign
- Hardware Trojan Threats to Multi-Chiplet Photonic Neural Network Accelerators
- Mapping Dynamic, Hierarchical Quantum Circuits
- Chiplet-Based Techniques for Scalable and Memory-Aware Multiscalar Multiplication on Hardware Platforms
- A Time-Encoded Analog Photonic Interposer for Energy-Efficient Integration of Analog Vision Sensors and Analog Accelerators