MGSim + MGMark: A Framework for Multi-GPU System Research

Sun, Yifan; Baruah, Trinayan; Mojumder, Saiful A.; Dong, Shi; Ubal, Rafael; Gong, Xiang; Treadway, Shane; Bao, Yuhui; Zhao, Vincent; Abellán, José L.; Kim, John; Joshi, Ajay; Kaeli, David

Computer Science > Distributed, Parallel, and Cluster Computing

arXiv:1811.02884 (cs)

[Submitted on 15 Oct 2018 (v1), last revised 13 Nov 2018 (this version, v3)]

Title:MGSim + MGMark: A Framework for Multi-GPU System Research

Authors:Yifan Sun, Trinayan Baruah, Saiful A. Mojumder, Shi Dong, Rafael Ubal, Xiang Gong, Shane Treadway, Yuhui Bao, Vincent Zhao, José L. Abellán, John Kim, Ajay Joshi, David Kaeli

View PDF

Abstract:The rapidly growing popularity and scale of data-parallel workloads demand a corresponding increase in raw computational power of GPUs (Graphics Processing Units). As single-GPU systems struggle to satisfy the performance demands, multi-GPU systems have begun to dominate the high-performance computing world. The advent of such systems raises a number of design challenges, including the GPU microarchitecture, multi-GPU interconnect fabrics, runtime libraries and associated programming models. The research community currently lacks a publically available and comprehensive multi-GPU simulation framework and benchmark suite to evaluate multi-GPU system design solutions.
In this work, we present MGSim, a cycle-accurate, extensively validated, multi-GPU simulator, based on AMD's Graphics Core Next 3 (GCN3) instruction set architecture. We complement MGSim with MGMark, a suite of multi-GPU workloads that explores multi-GPU collaborative execution patterns. Our simulator is scalable and comes with in-built support for multi-threaded execution to enable fast and efficient simulations. In terms of performance accuracy, MGSim differs $5.5\%$ on average when compared against actual GPU hardware. We also achieve a $3.5\times$ and a $2.5\times$ average speedup in function emulation and architectural simulation with 4 CPU cores, while delivering the same accuracy as the serial simulation.
We illustrate the novel simulation capabilities provided by our simulator through a case study exploring programming models based on a unified multi-GPU system (U-MGPU) and a discrete multi-GPU system (D-MGPU) that both utilize unified memory space and cross-GPU memory access. We evaluate the design implications from our case study, suggesting that D-MGPU is an attractive programming model for future multi-GPU systems.

Comments:	Updated typo
Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC); Hardware Architecture (cs.AR)
Cite as:	arXiv:1811.02884 [cs.DC]
	(or arXiv:1811.02884v3 [cs.DC] for this version)
	https://doi.org/10.48550/arXiv.1811.02884

Submission history

From: Yifan Sun [view email]
[v1] Mon, 15 Oct 2018 18:38:57 UTC (551 KB)
[v2] Thu, 8 Nov 2018 02:17:30 UTC (578 KB)
[v3] Tue, 13 Nov 2018 20:20:00 UTC (577 KB)

Computer Science > Distributed, Parallel, and Cluster Computing

Title:MGSim + MGMark: A Framework for Multi-GPU System Research

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Distributed, Parallel, and Cluster Computing

Title:MGSim + MGMark: A Framework for Multi-GPU System Research

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators