Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution

Gao, Yunquan; Zhang, Zhiguo; Donta, Praveen Kumar; Dehury, Chinmaya Kumar; Wang, Xiujun; Niyato, Dusit; Zhang, Qiyang

Computer Science > Distributed, Parallel, and Cluster Computing

arXiv:2503.21109 (cs)

[Submitted on 27 Mar 2025]

Title:Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution

Authors:Yunquan Gao, Zhiguo Zhang, Praveen Kumar Donta, Chinmaya Kumar Dehury, Xiujun Wang, Dusit Niyato, Qiyang Zhang

View PDF HTML (experimental)

Abstract:Deep Neural Networks (DNNs) are increasingly deployed across diverse industries, driving demand for mobile device support. However, existing mobile inference frameworks often rely on a single processor per model, limiting hardware utilization and causing suboptimal performance and energy efficiency. Expanding DNN accessibility on mobile platforms requires adaptive, resource-efficient solutions to meet rising computational needs without compromising functionality. Parallel inference of multiple DNNs on heterogeneous processors remains challenging. Some works partition DNN operations into subgraphs for parallel execution across processors, but these often create excessive subgraphs based only on hardware compatibility, increasing scheduling complexity and memory overhead.
To address this, we propose an Advanced Multi-DNN Model Scheduling (ADMS) strategy for optimizing multi-DNN inference on mobile heterogeneous processors. ADMS constructs an optimal subgraph partitioning strategy offline, balancing hardware operation support and scheduling granularity, and uses a processor-state-aware algorithm to dynamically adjust workloads based on real-time conditions. This ensures efficient workload distribution and maximizes processor utilization. Experiments show ADMS reduces multi-DNN inference latency by 4.04 times compared to vanilla frameworks.

Comments:	14 pages, 12 figures, 5 tables
Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
MSC classes:	68T07, 68W40
ACM classes:	I.2.6; C.1.4; D.4.8
Cite as:	arXiv:2503.21109 [cs.DC]
	(or arXiv:2503.21109v1 [cs.DC] for this version)
	https://doi.org/10.48550/arXiv.2503.21109

Submission history

From: Zhiguo Zhang [view email]
[v1] Thu, 27 Mar 2025 03:03:09 UTC (25,952 KB)

Computer Science > Distributed, Parallel, and Cluster Computing

Title:Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Distributed, Parallel, and Cluster Computing

Title:Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators