A performance comparison of Dask and Apache Spark for data-intensive neuroimaging pipelines

Dugré, Mathieu; Hayot-Sasson, Valérie; Glatard, Tristan

Computer Science > Distributed, Parallel, and Cluster Computing

arXiv:1907.13030 (cs)

[Submitted on 30 Jul 2019 (v1), last revised 5 Oct 2019 (this version, v3)]

Title:A performance comparison of Dask and Apache Spark for data-intensive neuroimaging pipelines

Authors:Mathieu Dugré, Valérie Hayot-Sasson, Tristan Glatard

View PDF

Abstract:In the past few years, neuroimaging has entered the Big Data era due to the joint increase in image resolution, data sharing, and study sizes. However, no particular Big Data engines have emerged in this field, and several alternatives remain available. We compare two popular Big Data engines with Python APIs, Apache Spark and Dask, for their runtime performance in processing neuroimaging pipelines. Our evaluation uses two synthetic pipelines processing the 81GB BigBrain image, and a real pipeline processing anatomical data from more than 1,000 subjects. We benchmark these pipelines using various combinations of task durations, data sizes, and numbers of workers, deployed on an 8-node (8 cores ea.) compute cluster in Compute Canada's Arbutus cloud. We evaluate PySpark's RDD API against Dask's Bag, Delayed and Futures. Results show that despite slight differences between Spark and Dask, both engines perform comparably. However, Dask pipelines risk being limited by Python's GIL depending on task type and cluster configuration. In all cases, the major limiting factor was data transfer. While either engine is suitable for neuroimaging pipelines, more effort needs to be placed in reducing data transfer time.

Comments:	10 pages, 15 figures, 1 tables. To appear in the proceeding of the 14th WORKS Workshop on Topics in Workflows in Support of Large-Scale Science, 17 November 2019, Denver, CO, USA
Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC); Performance (cs.PF)
Cite as:	arXiv:1907.13030 [cs.DC]
	(or arXiv:1907.13030v3 [cs.DC] for this version)
	https://doi.org/10.48550/arXiv.1907.13030

Submission history

From: Mathieu Dugré [view email]
[v1] Tue, 30 Jul 2019 15:42:32 UTC (2,828 KB)
[v2] Wed, 31 Jul 2019 19:38:32 UTC (2,828 KB)
[v3] Sat, 5 Oct 2019 20:24:36 UTC (2,837 KB)

Computer Science > Distributed, Parallel, and Cluster Computing

Title:A performance comparison of Dask and Apache Spark for data-intensive neuroimaging pipelines

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Distributed, Parallel, and Cluster Computing

Title:A performance comparison of Dask and Apache Spark for data-intensive neuroimaging pipelines

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators