Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection

Fan, Cunhang; Ding, Mingming; Tao, Jianhua; Fu, Ruibo; Yi, Jiangyan; Wen, Zhengqi; Lv, Zhao

doi:10.1109/TASLP.2024.3389643

Computer Science > Sound

arXiv:2310.08869 (cs)

[Submitted on 13 Oct 2023 (v1), last revised 16 Apr 2024 (this version, v2)]

Title:Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection

Authors:Cunhang Fan, Mingming Ding, Jianhua Tao, Ruibo Fu, Jiangyan Yi, Zhengqi Wen, Zhao Lv

View PDF HTML (experimental)

Abstract:Most research in synthetic speech detection (SSD) focuses on improving performance on standard noise-free datasets. However, in actual situations, noise interference is usually present, causing significant performance degradation in SSD systems. To improve noise robustness, this paper proposes a dual-branch knowledge distillation synthetic speech detection (DKDSSD) method. Specifically, a parallel data flow of the clean teacher branch and the noisy student branch is designed, and interactive fusion module and response-based teacher-student paradigms are proposed to guide the training of noisy data from both the data distribution and decision-making perspectives. In the noisy student branch, speech enhancement is introduced initially for denoising, aiming to reduce the interference of strong noise. The proposed interactive fusion combines denoised features and noisy features to mitigate the impact of speech distortion and ensure consistency with the data distribution of the clean branch. The teacher-student paradigm maps the student's decision space to the teacher's decision space, enabling noisy speech to behave similarly to clean speech. Additionally, a joint training method is employed to optimize both branches for achieving global optimality. Experimental results based on multiple datasets demonstrate that the proposed method performs effectively in noisy environments and maintains its performance in cross-dataset experiments. Source code is available at this https URL.

Subjects:	Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2310.08869 [cs.SD]
	(or arXiv:2310.08869v2 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.2310.08869
Related DOI:	https://doi.org/10.1109/TASLP.2024.3389643

Submission history

From: Mingming Ding [view email]
[v1] Fri, 13 Oct 2023 05:37:29 UTC (14,868 KB)
[v2] Tue, 16 Apr 2024 10:18:50 UTC (5,956 KB)

Computer Science > Sound

Title:Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators