Audio and Speech Processing

Authors and titles for April 2021

Total of 266 entries : 1-25 ... 151-175 176-200 201-225 226-250 251-266

Showing up to 25 entries per page: fewer | more | all

[226] arXiv:2104.10507 (cross-list from cs.CL) [pdf, other]: Title: On Sampling-Based Training Criteria for Neural Language Modeling

Yingbo Gao, David Thulke, Alexander Gerstenberger, Khoa Viet Tran, Ralf Schlüter, Hermann Ney

Comments: Accepted at INTERSPEECH 2021

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
[227] arXiv:2104.10747 (cross-list from cs.CL) [pdf, other]: Title: Accented Speech Recognition: A Survey

Arthur Hinsvark (1), Natalie Delworth (1), Miguel Del Rio (1), Quinten McNamara (1), Joshua Dong (1), Ryan Westerman (1), Michelle Huang (1), Joseph Palakapilly (1), Jennifer Drexler (1), Ilya Pirkin (1), Nishchal Bhandari (1), Miguel Jette (1) ((1) <a href="http://Rev.com" rel="external noopener nofollow" class="link-external link-http">this http URL</a>)

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[228] arXiv:2104.11051 (cross-list from cs.SD) [pdf, other]: Title: Protecting gender and identity with disentangled speech representations

Dimitrios Stoidis, Andrea Cavallaro

Comments: 5 pages, 2 figures

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[229] arXiv:2104.11116 (cross-list from cs.CV) [pdf, other]: Title: Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation

Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, Ziwei Liu

Comments: Accepted to IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. Code and models are available at this https URL

Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[230] arXiv:2104.11127 (cross-list from cs.CL) [pdf, other]: Title: Fast Text-Only Domain Adaptation of RNN-Transducer Prediction Network

Janne Pylkkönen (1), Antti Ukkonen (1 and 2), Juho Kilpikoski (1), Samu Tamminen (1), Hannes Heikinheimo (1) ((1) Speechly, (2) Department of Computer Science, University of Helsinki, Finland)

Comments: 5 pages, 2 figures. Accepted to Interspeech 2021

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[231] arXiv:2104.11347 (cross-list from cs.SD) [pdf, other]: Title: Restoring degraded speech via a modified diffusion model

Jianwei Zhang, Suren Jayasuriya, Visar Berisha

Journal-ref: Proc. Interspeech 2021, 221-225, 2021)

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[232] arXiv:2104.11348 (cross-list from cs.CL) [pdf, other]: Title: Earnings-21: A Practical Benchmark for ASR in the Wild

Miguel Del Rio, Natalie Delworth, Ryan Westerman, Michelle Huang, Nishchal Bhandari, Joseph Palakapilly, Quinten McNamara, Joshua Dong, Piotr Zelasko, Miguel Jette

Comments: Accepted to INTERSPEECH 2021. June 15 2021: Addressing the comments of reviewers and updating the results of our internal ESPNet model. The results do not change our conclusions. April 28th, 2021: We found and resolved an issue in our experimental evaluation that scored the LibriSpeech model at ~20% worse relative WER than the actual WER. The updated results do not affect our conclusions

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[233] arXiv:2104.11395 (cross-list from cs.SD) [pdf, other]: Title: Infant Vocal Tract Development Analysis and Diagnosis by Cry Signals with CNN Age Classification

Chunyan Ji, Yi Pan

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[234] arXiv:2104.11462 (cross-list from cs.CL) [pdf, other]: Title: LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from Speech

Solene Evain, Ha Nguyen, Hang Le, Marcely Zanon Boito, Salima Mdhaffar, Sina Alisamir, Ziyi Tong, Natalia Tomashenko, Marco Dinarelli, Titouan Parcollet, Alexandre Allauzen, Yannick Esteve, Benjamin Lecouteux, Francois Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier

Comments: Will be presented at Interspeech 2021

Journal-ref: Proc. Interspeech 2021

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[235] arXiv:2104.11532 (cross-list from cs.SD) [pdf, other]: Title: 3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces

László Tóth, Amin Honarmandi Shandiz

Comments: 10 pages, 2 tables , 3 figures

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[236] arXiv:2104.11587 (cross-list from cs.SD) [pdf, other]: Title: ESResNe(X)t-fbsp: Learning Robust Time-Frequency Transformation of Audio

Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel

Comments: submitted IJCNN 2021

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[237] arXiv:2104.11598 (cross-list from cs.SD) [pdf, other]: Title: Reconstructing Speech from Real-Time Articulatory MRI Using Neural Vocoders

Yide Yu, Amin Honarmandi Shandiz, László Tóth

Comments: 6 pages. 4 tables, 3 figures

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[238] arXiv:2104.11601 (cross-list from cs.SD) [pdf, other]: Title: Improving Neural Silent Speech Interface Models by Adversarial Training

Amin Honarmandi Shandiz, László Tóth, Gábor Gosztolya, Alexandra Markó, Tamás Gábor Csapó

Comments: 11 pages, 3 tables, 2 figures

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[239] arXiv:2104.11629 (cross-list from cs.SD) [pdf, other]: Title: DeepSpectrumLite: A Power-Efficient Transfer Learning Framework for Embedded Speech and Audio Processing from Decentralised Data

Shahin Amiriparian (1), Tobias Hübner (1), Maurice Gerczuk (1), Sandra Ottl (1), Björn W. Schuller (1,2) ((1) EIHW -- Chair of Embedded Intelligence for Health Care and Wellbeing, University of Augsburg, Germany, (2) GLAM -- Group on Language, Audio, and Music, Imperial College London, UK)

Comments: 5 pages, 1 figure

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[240] arXiv:2104.11673 (cross-list from cs.SD) [pdf, other]: Title: Deep Learning Based Assessment of Synthetic Speech Naturalness

Gabriel Mittag, Sebastian Möller

Comments: Late upload, presented at Interspeech 2020

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[241] arXiv:2104.11710 (cross-list from cs.SD) [pdf, other]: Title: Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation

Marco Gaido, Matteo Negri, Mauro Cettolo, Marco Turchi

Comments: Accepted to ICNLSP 2021

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[242] arXiv:2104.11880 (cross-list from cs.SD) [pdf, other]: Title: Music Embedding: A Tool for Incorporating Music Theory into Computational Music Applications

SeyyedPooya HekmatiAthar, Mohd Anwar

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[243] arXiv:2104.11946 (cross-list from cs.LG) [pdf, other]: Title: Aligned Contrastive Predictive Coding

Jan Chorowski, Grzegorz Ciesielski, Jarosław Dzikowski, Adrian Łańcucki, Ricard Marxer, Mateusz Opala, Piotr Pusz, Paweł Rychlikowski, Michał Stypułkowski

Comments: Published in Interspeech 2021

Subjects: Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[244] arXiv:2104.11984 (cross-list from cs.SD) [pdf, other]: Title: MusCaps: Generating Captions for Music Audio

Ilaria Manco, Emmanouil Benetos, Elio Quinton, Gyorgy Fazekas

Comments: Accepted to IJCNN 2021 for the Special Session on Representation Learning for Audio, Speech, and Music Processing

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[245] arXiv:2104.12159 (cross-list from cs.SD) [pdf, other]: Title: An Adaptive Learning based Generative Adversarial Network for One-To-One Voice Conversion

Sandipan Dhar, Nanda Dulal Jana, Swagatam Das

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[246] arXiv:2104.12292 (cross-list from cs.SD) [pdf, other]: Title: Text-to-Speech Synthesis Techniques for MIDI-to-Audio Synthesis

Erica Cooper, Xin Wang, Junichi Yamagishi

Comments: In the proceedings of ISCA Speech Synthesis Workshop 2021

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[247] arXiv:2104.12359 (cross-list from cs.SD) [pdf, other]: Title: Complex Neural Spatial Filter: Enhancing Multi-channel Target Speech Separation in Complex Domain

Rongzhi Gu, Shi-Xiong Zhang, Yuexian Zou, Dong Yu

Comments: 5 pages, 3 figures

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[248] arXiv:2104.12432 (cross-list from cs.SD) [pdf, other]: Title: Generation of musical patterns through operads

Samuele Giraudo

Comments: 10 pages

Journal-ref: Journ\'ees d'informatique musicale, 2020

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Combinatorics (math.CO)
[249] arXiv:2104.12462 (cross-list from cs.SD) [pdf, other]: Title: Points2Sound: From mono to binaural audio using 3D point cloud scenes

Francesc Lluís, Vasileios Chatziioannou, Alex Hofmann

Comments: Code, data, and listening examples: this https URL

Journal-ref: EURASIP Journal on Audio, Speech, and Music Processing 2022 (1), 1-15

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[250] arXiv:2104.12693 (cross-list from cs.SD) [pdf, other]: Title: Identifying Actions for Sound Event Classification

Benjamin Elizalde, Radu Revutchi, Samarjit Das, Bhiksha Raj, Ian Lane, Laurie M. Heller

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)

Total of 266 entries : 1-25 ... 151-175 176-200 201-225 226-250 251-266

Showing up to 25 entries per page: fewer | more | all