Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems

Joshi, Sonal; Villalba, Jesús; Żelasko, Piotr; Moro-Velázquez, Laureano; Dehak, Najim

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2101.08909 (eess)

[Submitted on 22 Jan 2021 (v1), last revised 25 Jun 2021 (this version, v2)]

Title:Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems

Authors:Sonal Joshi, Jesús Villalba, Piotr Żelasko, Laureano Moro-Velázquez, Najim Dehak

View PDF

Abstract:Adversarial examples to speaker recognition (SR) systems are generated by adding a carefully crafted noise to the speech signal to make the system fail while being imperceptible to humans. Such attacks pose severe security risks, making it vital to deep-dive and understand how much the state-of-the-art SR systems are vulnerable to these attacks. Moreover, it is of greater importance to propose defenses that can protect the systems against these attacks. Addressing these concerns, this paper at first investigates how state-of-the-art x-vector based SR systems are affected by white-box adversarial attacks, i.e., when the adversary has full knowledge of the system. x-Vector based SR systems are evaluated against white-box adversarial attacks common in the literature like fast gradient sign method (FGSM), basic iterative method (BIM)--a.k.a. iterative-FGSM--, projected gradient descent (PGD), and Carlini-Wagner (CW) attack. To mitigate against these attacks, the paper proposes four pre-processing defenses. It evaluates them against powerful adaptive white-box adversarial attacks, i.e., when the adversary has full knowledge of the system, including the defense. The four pre-processing defenses--viz. randomized smoothing, DefenseGAN, variational autoencoder (VAE), and Parallel WaveGAN vocoder (PWG) are compared against the baseline defense of adversarial training. Conclusions indicate that SR systems were extremely vulnerable under BIM, PGD, and CW attacks. Among the proposed pre-processing defenses, PWG combined with randomized smoothing offers the most protection against the attacks, with accuracy averaging 93% compared to 52% in the undefended system and an absolute improvement >90% for BIM attacks with $L_\infty>0.001$ and CW attack.

Comments:	This work has been submitted to the IEEE for possible publication
Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as:	arXiv:2101.08909 [eess.AS]
	(or arXiv:2101.08909v2 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2101.08909

Submission history

From: Sonal Joshi [view email]
[v1] Fri, 22 Jan 2021 01:25:18 UTC (1,014 KB)
[v2] Fri, 25 Jun 2021 15:37:47 UTC (1,860 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators