Analysis of ABC Frontend Audio Systems for the NIST-SRE24
Authors:
Sara Barahona,
Anna Silnova,
Ladislav Mošner,
Junyi Peng,
Oldřich Plchot,
Johan Rohdin,
Lin Zhang,
Jiangyu Han,
Petr Palka,
Federico Landini,
Lukáš Burget,
Themos Stafylakis,
Sandro Cumani,
Dominik Boboš,
Miroslav Hlavaček,
Martin Kodovsky,
Tomáš Pavlíček
Abstract:
We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by NIST: using only a provided set of telephone recordings for training (fixed) or adding publicly available data (open condition). Under these constraints, we develop the best possible speaker embedding extractors for the p…
▽ More
We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by NIST: using only a provided set of telephone recordings for training (fixed) or adding publicly available data (open condition). Under these constraints, we develop the best possible speaker embedding extractors for the pre-dominant conversational telephone speech (CTS) domain. We explored architectures based on ResNet with different pooling mechanisms, recently introduced ReDimNet architecture, as well as a system based on the XLS-R model, which represents the family of large pre-trained self-supervised models. In open condition, we train on VoxBlink2 dataset, containing 110 thousand speakers across multiple languages. We observed a good performance and robustness of VoxBlink-trained models, and our experiments show practical recipes for developing state-of-the-art frontends for speaker recognition.
△ Less
Submitted 21 May, 2025;
originally announced May 2025.
State-of-the-art Embeddings with Video-free Segmentation of the Source VoxCeleb Data
Authors:
Sara Barahona,
Ladislav Mošner,
Themos Stafylakis,
Oldřich Plchot,
Junyi Peng,
Lukáš Burget,
Jan Černocký
Abstract:
In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that…
▽ More
In this paper, we refine and validate our method for training speaker embedding extractors using weak annotations. More specifically, we use only the audio stream of the source VoxCeleb videos and the names of the celebrities without knowing the time intervals in which they appear in the recording. We experiment with hyperparameters and embedding extractors based on ResNet and WavLM. We show that the method achieves state-of-the-art results in speaker verification, comparable with training the extractors in a standard supervised way on the VoxCeleb dataset. We also extend it by considering segments belonging to unknown speakers appearing alongside the celebrities, which are typically being discarded. Overall, our approach can be used for directly training state-of-the-art embedding extractors or as an alternative to the VoxCeleb-like pipeline for dataset creation without needing image modality.
△ Less
Submitted 3 October, 2024;
originally announced October 2024.