Skip to main content
Cornell University
We gratefully acknowledge support from the Simons Foundation, member institutions, and all contributors. Donate
arxiv logo > eess.AS

Help | Advanced Search

arXiv logo
Cornell University Logo

quick links

  • Login
  • Help Pages
  • About

Audio and Speech Processing

Authors and titles for April 2021

Total of 266 entries : 1-25 ... 151-175 176-200 201-225 226-250 251-266
Showing up to 25 entries per page: fewer | more | all
[226] arXiv:2104.10507 (cross-list from cs.CL) [pdf, other]
Title: On Sampling-Based Training Criteria for Neural Language Modeling
Yingbo Gao, David Thulke, Alexander Gerstenberger, Khoa Viet Tran, Ralf Schlüter, Hermann Ney
Comments: Accepted at INTERSPEECH 2021
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
[227] arXiv:2104.10747 (cross-list from cs.CL) [pdf, other]
Title: Accented Speech Recognition: A Survey
Arthur Hinsvark (1), Natalie Delworth (1), Miguel Del Rio (1), Quinten McNamara (1), Joshua Dong (1), Ryan Westerman (1), Michelle Huang (1), Joseph Palakapilly (1), Jennifer Drexler (1), Ilya Pirkin (1), Nishchal Bhandari (1), Miguel Jette (1) ((1) <a href="http://Rev.com" rel="external noopener nofollow" class="link-external link-http">this http URL</a>)
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[228] arXiv:2104.11051 (cross-list from cs.SD) [pdf, other]
Title: Protecting gender and identity with disentangled speech representations
Dimitrios Stoidis, Andrea Cavallaro
Comments: 5 pages, 2 figures
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[229] arXiv:2104.11116 (cross-list from cs.CV) [pdf, other]
Title: Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, Ziwei Liu
Comments: Accepted to IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. Code and models are available at this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[230] arXiv:2104.11127 (cross-list from cs.CL) [pdf, other]
Title: Fast Text-Only Domain Adaptation of RNN-Transducer Prediction Network
Janne Pylkkönen (1), Antti Ukkonen (1 and 2), Juho Kilpikoski (1), Samu Tamminen (1), Hannes Heikinheimo (1) ((1) Speechly, (2) Department of Computer Science, University of Helsinki, Finland)
Comments: 5 pages, 2 figures. Accepted to Interspeech 2021
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[231] arXiv:2104.11347 (cross-list from cs.SD) [pdf, other]
Title: Restoring degraded speech via a modified diffusion model
Jianwei Zhang, Suren Jayasuriya, Visar Berisha
Journal-ref: Proc. Interspeech 2021, 221-225, 2021)
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[232] arXiv:2104.11348 (cross-list from cs.CL) [pdf, other]
Title: Earnings-21: A Practical Benchmark for ASR in the Wild
Miguel Del Rio, Natalie Delworth, Ryan Westerman, Michelle Huang, Nishchal Bhandari, Joseph Palakapilly, Quinten McNamara, Joshua Dong, Piotr Zelasko, Miguel Jette
Comments: Accepted to INTERSPEECH 2021. June 15 2021: Addressing the comments of reviewers and updating the results of our internal ESPNet model. The results do not change our conclusions. April 28th, 2021: We found and resolved an issue in our experimental evaluation that scored the LibriSpeech model at ~20% worse relative WER than the actual WER. The updated results do not affect our conclusions
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[233] arXiv:2104.11395 (cross-list from cs.SD) [pdf, other]
Title: Infant Vocal Tract Development Analysis and Diagnosis by Cry Signals with CNN Age Classification
Chunyan Ji, Yi Pan
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[234] arXiv:2104.11462 (cross-list from cs.CL) [pdf, other]
Title: LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from Speech
Solene Evain, Ha Nguyen, Hang Le, Marcely Zanon Boito, Salima Mdhaffar, Sina Alisamir, Ziyi Tong, Natalia Tomashenko, Marco Dinarelli, Titouan Parcollet, Alexandre Allauzen, Yannick Esteve, Benjamin Lecouteux, Francois Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier
Comments: Will be presented at Interspeech 2021
Journal-ref: Proc. Interspeech 2021
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[235] arXiv:2104.11532 (cross-list from cs.SD) [pdf, other]
Title: 3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces
László Tóth, Amin Honarmandi Shandiz
Comments: 10 pages, 2 tables , 3 figures
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[236] arXiv:2104.11587 (cross-list from cs.SD) [pdf, other]
Title: ESResNe(X)t-fbsp: Learning Robust Time-Frequency Transformation of Audio
Andrey Guzhov, Federico Raue, Jörn Hees, Andreas Dengel
Comments: submitted IJCNN 2021
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[237] arXiv:2104.11598 (cross-list from cs.SD) [pdf, other]
Title: Reconstructing Speech from Real-Time Articulatory MRI Using Neural Vocoders
Yide Yu, Amin Honarmandi Shandiz, László Tóth
Comments: 6 pages. 4 tables, 3 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[238] arXiv:2104.11601 (cross-list from cs.SD) [pdf, other]
Title: Improving Neural Silent Speech Interface Models by Adversarial Training
Amin Honarmandi Shandiz, László Tóth, Gábor Gosztolya, Alexandra Markó, Tamás Gábor Csapó
Comments: 11 pages, 3 tables, 2 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[239] arXiv:2104.11629 (cross-list from cs.SD) [pdf, other]
Title: DeepSpectrumLite: A Power-Efficient Transfer Learning Framework for Embedded Speech and Audio Processing from Decentralised Data
Shahin Amiriparian (1), Tobias Hübner (1), Maurice Gerczuk (1), Sandra Ottl (1), Björn W. Schuller (1,2) ((1) EIHW -- Chair of Embedded Intelligence for Health Care and Wellbeing, University of Augsburg, Germany, (2) GLAM -- Group on Language, Audio, and Music, Imperial College London, UK)
Comments: 5 pages, 1 figure
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[240] arXiv:2104.11673 (cross-list from cs.SD) [pdf, other]
Title: Deep Learning Based Assessment of Synthetic Speech Naturalness
Gabriel Mittag, Sebastian Möller
Comments: Late upload, presented at Interspeech 2020
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[241] arXiv:2104.11710 (cross-list from cs.SD) [pdf, other]
Title: Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation
Marco Gaido, Matteo Negri, Mauro Cettolo, Marco Turchi
Comments: Accepted to ICNLSP 2021
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[242] arXiv:2104.11880 (cross-list from cs.SD) [pdf, other]
Title: Music Embedding: A Tool for Incorporating Music Theory into Computational Music Applications
SeyyedPooya HekmatiAthar, Mohd Anwar
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[243] arXiv:2104.11946 (cross-list from cs.LG) [pdf, other]
Title: Aligned Contrastive Predictive Coding
Jan Chorowski, Grzegorz Ciesielski, Jarosław Dzikowski, Adrian Łańcucki, Ricard Marxer, Mateusz Opala, Piotr Pusz, Paweł Rychlikowski, Michał Stypułkowski
Comments: Published in Interspeech 2021
Subjects: Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[244] arXiv:2104.11984 (cross-list from cs.SD) [pdf, other]
Title: MusCaps: Generating Captions for Music Audio
Ilaria Manco, Emmanouil Benetos, Elio Quinton, Gyorgy Fazekas
Comments: Accepted to IJCNN 2021 for the Special Session on Representation Learning for Audio, Speech, and Music Processing
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[245] arXiv:2104.12159 (cross-list from cs.SD) [pdf, other]
Title: An Adaptive Learning based Generative Adversarial Network for One-To-One Voice Conversion
Sandipan Dhar, Nanda Dulal Jana, Swagatam Das
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[246] arXiv:2104.12292 (cross-list from cs.SD) [pdf, other]
Title: Text-to-Speech Synthesis Techniques for MIDI-to-Audio Synthesis
Erica Cooper, Xin Wang, Junichi Yamagishi
Comments: In the proceedings of ISCA Speech Synthesis Workshop 2021
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[247] arXiv:2104.12359 (cross-list from cs.SD) [pdf, other]
Title: Complex Neural Spatial Filter: Enhancing Multi-channel Target Speech Separation in Complex Domain
Rongzhi Gu, Shi-Xiong Zhang, Yuexian Zou, Dong Yu
Comments: 5 pages, 3 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[248] arXiv:2104.12432 (cross-list from cs.SD) [pdf, other]
Title: Generation of musical patterns through operads
Samuele Giraudo
Comments: 10 pages
Journal-ref: Journ\'ees d'informatique musicale, 2020
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Combinatorics (math.CO)
[249] arXiv:2104.12462 (cross-list from cs.SD) [pdf, other]
Title: Points2Sound: From mono to binaural audio using 3D point cloud scenes
Francesc Lluís, Vasileios Chatziioannou, Alex Hofmann
Comments: Code, data, and listening examples: this https URL
Journal-ref: EURASIP Journal on Audio, Speech, and Music Processing 2022 (1), 1-15
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[250] arXiv:2104.12693 (cross-list from cs.SD) [pdf, other]
Title: Identifying Actions for Sound Event Classification
Benjamin Elizalde, Radu Revutchi, Samarjit Das, Bhiksha Raj, Ian Lane, Laurie M. Heller
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
Total of 266 entries : 1-25 ... 151-175 176-200 201-225 226-250 251-266
Showing up to 25 entries per page: fewer | more | all
  • About
  • Help
  • contact arXivClick here to contact arXiv Contact
  • subscribe to arXiv mailingsClick here to subscribe Subscribe
  • Copyright
  • Privacy Policy
  • Web Accessibility Assistance
  • arXiv Operational Status
    Get status notifications via email or slack