FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation

Chen, Jianyi; Xue, Wei; Tan, Xu; Ye, Zhen; Liu, Qifeng; Guo, Yike

Computer Science > Sound

arXiv:2405.07682 (cs)

[Submitted on 13 May 2024]

Title:FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation

Authors:Jianyi Chen, Wei Xue, Xu Tan, Zhen Ye, Qifeng Liu, Yike Guo

View PDF HTML (experimental)

Abstract:Singing Accompaniment Generation (SAG), which generates instrumental music to accompany input vocals, is crucial to developing human-AI symbiotic art creation systems. The state-of-the-art method, SingSong, utilizes a multi-stage autoregressive (AR) model for SAG, however, this method is extremely slow as it generates semantic and acoustic tokens recursively, and this makes it impossible for real-time applications. In this paper, we aim to develop a Fast SAG method that can create high-quality and coherent accompaniments. A non-AR diffusion-based framework is developed, which by carefully designing the conditions inferred from the vocal signals, generates the Mel spectrogram of the target accompaniment directly. With diffusion and Mel spectrogram modeling, the proposed method significantly simplifies the AR token-based SingSong framework, and largely accelerates the generation. We also design semantic projection, prior projection blocks as well as a set of loss functions, to ensure the generated accompaniment has semantic and rhythm coherence with the vocal signal. By intensive experimental studies, we demonstrate that the proposed method can generate better samples than SingSong, and accelerate the generation by at least 30 times. Audio samples and code are available at this https URL.

Comments:	IJCAI 2024
Subjects:	Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2405.07682 [cs.SD]
	(or arXiv:2405.07682v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2405.07682

Submission history

From: Wei Xue [view email]
[v1] Mon, 13 May 2024 12:14:54 UTC (1,009 KB)

Computer Science > Sound

Title:FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators