DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

Wang, Qing; Yao, Jixun; Sun, Zhaokai; Guo, Pengcheng; Xie, Lei; Hansen, John H. L.

Computer Science > Sound

arXiv:2501.05127 (cs)

[Submitted on 9 Jan 2025]

Title:DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

Authors:Qing Wang, Jixun Yao, Zhaokai Sun, Pengcheng Guo, Lei Xie, John H.L. Hansen

View PDF HTML (experimental)

Abstract:Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we aim to perform more realistic attacks in SID, which are challenging for both humans and machines to detect. In this study, we propose DiffAttack, a novel timbre-reserved adversarial attack approach that exploits the capability of a diffusion-based voice conversion (DiffVC) model to generate adversarial fake audio with distinct target speaker attribution. By introducing adversarial constraints into the generative process of the diffusion-based voice conversion model, we craft fake samples that effectively mislead target models while preserving speaker-wise characteristics. Specifically, inspired by the use of randomly sampled Gaussian noise in conventional adversarial attacks and diffusion processes, we incorporate adversarial constraints into the reverse diffusion process. These constraints subtly guide the reverse diffusion process toward aligning with the target speaker distribution. Our experiments on the LibriTTS dataset indicate that DiffAttack significantly improves the attack success rate compared to vanilla DiffVC and other methods. Moreover, objective and subjective evaluations demonstrate that introducing adversarial constraints does not compromise the speech quality generated by the DiffVC model.

Comments:	5 pages,4 figures, accepted by ICASSP 2025
Subjects:	Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2501.05127 [cs.SD]
	(or arXiv:2501.05127v1 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2501.05127

Submission history

From: Qing Wang [view email]
[v1] Thu, 9 Jan 2025 10:30:58 UTC (2,320 KB)

Computer Science > Sound

Title:DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators