Self Multi-Head Attention for Speaker Recognition

India, Miquel; Safari, Pooyan; Hernando, Javier

Computer Science > Sound

arXiv:1906.09890 (cs)

[Submitted on 24 Jun 2019 (v1), last revised 1 Jul 2019 (this version, v2)]

Title:Self Multi-Head Attention for Speaker Recognition

Authors:Miquel India, Pooyan Safari, Javier Hernando

View PDF

Abstract:Most state-of-the-art Deep Learning (DL) approaches for speaker recognition work on a short utterance level. Given the speech signal, these algorithms extract a sequence of speaker embeddings from short segments and those are averaged to obtain an utterance level speaker representation. In this work we propose the use of an attention mechanism to obtain a discriminative speaker embedding given non fixed length speech utterances. Our system is based on a Convolutional Neural Network (CNN) that encodes short-term speaker features from the spectrogram and a self multi-head attention model that maps these representations into a long-term speaker embedding. The attention model that we propose produces multiple alignments from different subsegments of the CNN encoded states over the sequence. Hence this mechanism works as a pooling layer which decides the most discriminative features over the sequence to obtain an utterance level representation. We have tested this approach for the verification task for the VoxCeleb1 dataset. The results show that self multi-head attention outperforms both temporal and statistical pooling methods with a 18\% of relative EER. Obtained results show a 58\% relative improvement in EER compared to i-vector+PLDA.

Comments:	4+1 pages. 4 Figures. Accepted for Interspeech 2009
Subjects:	Sound (cs.SD); Machine Learning (cs.LG); Machine Learning (stat.ML)
MSC classes:	68
Cite as:	arXiv:1906.09890 [cs.SD]
	(or arXiv:1906.09890v2 [cs.SD] for this version)
	https://doi.org/10.48550/arXiv.1906.09890

Submission history

From: Miquel India [view email]
[v1] Mon, 24 Jun 2019 12:44:09 UTC (594 KB)
[v2] Mon, 1 Jul 2019 22:02:09 UTC (594 KB)

Computer Science > Sound

Title:Self Multi-Head Attention for Speaker Recognition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Sound

Title:Self Multi-Head Attention for Speaker Recognition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators