Skip to main content
Cornell University
We gratefully acknowledge support from the Simons Foundation, member institutions, and all contributors. Donate
arxiv logo > eess.AS

Help | Advanced Search

arXiv logo
Cornell University Logo

quick links

  • Login
  • Help Pages
  • About

Audio and Speech Processing

Authors and titles for February 2021

Total of 208 entries : 1-50 51-100 101-150 126-175 151-200 201-208
Showing up to 50 entries per page: fewer | more | all
[126] arXiv:2102.04488 (cross-list from cs.CL) [pdf, other]
Title: Wake Word Detection with Streaming Transformers
Yiming Wang, Hang Lv, Daniel Povey, Lei Xie, Sanjeev Khudanpur
Comments: Accepted at IEEE ICASSP 2021. 5 pages, 3 figures
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[127] arXiv:2102.04588 (cross-list from cs.SD) [pdf, other]
Title: A comparative study of two-dimensional vocal tract acoustic modeling based on Finite-Difference Time-Domain methods
Debasish Ray Mohapatra, Victor Zappi, Sidney Fels
Comments: 4 pages, 3 figures
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[128] arXiv:2102.04680 (cross-list from cs.SD) [pdf, other]
Title: TräumerAI: Dreaming Music with StyleGAN
Dasaem Jeong, Seungheon Doh, Taegyun Kwon
Comments: presented in NeurIPS Workshop 2020: Machine Learning for Creativity and Design
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[129] arXiv:2102.04740 (cross-list from stat.ME) [pdf, other]
Title: Principal components variable importance reconstruction (PC-VIR): Exploring predictive importance in multicollinear acoustic speech data
Christopher Carignan, Ander Egurtzegi
Comments: 10 pages, 3 figures, GitHub repository
Subjects: Methodology (stat.ME); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[130] arXiv:2102.04832 (cross-list from eess.SP) [pdf, other]
Title: Fast and Accurate Amplitude Demodulation of Wideband Signals
Mantas Gabrielaitis
Comments: Accepted for publication in IEEE Transactions on Signal Processing
Subjects: Signal Processing (eess.SP); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[131] arXiv:2102.04880 (cross-list from cs.SD) [pdf, other]
Title: Diagnosis of COVID-19 and Non-COVID-19 Patients by Classifying Only a Single Cough Sound
Masoud Maleki
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Optimization and Control (math.OC)
[132] arXiv:2102.04932 (cross-list from cs.LG) [pdf, other]
Title: Sparsification via Compressed Sensing for Automatic Speech Recognition
Kai Zhen (1 and 2), Hieu Duy Nguyen (2), Feng-Ju Chang (2), Athanasios Mouchtaris (2), Ariya Rastrow (2). ((1) Indiana University Bloomington, (2) Alexa Machine Learning, Amazon, USA)
Comments: 5 pages, accepted for publication in (ICASSP 2021) 2021 IEEE International Conference on Acoustics, Speech, and Signal Processing. June 6-12, 2021. Location: Toronto, ON, Canada
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[133] arXiv:2102.04945 (cross-list from cs.SD) [pdf, other]
Title: On permutation invariant training for speech source separation
Xiaoyu Liu, Jordi Pons
Comments: In proceedings of ICASSP2021
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[134] arXiv:2102.04997 (cross-list from cs.LG) [pdf, other]
Title: Deep Neural Network based Cough Detection using Bed-mounted Accelerometer Measurements
Madhurananda Pahar, Igor Miranda, Andreas Diacon, Thomas Niesler
Comments: It has been accepted in ICASSP, 2021. Copyright information is shown at the very first page
Journal-ref: ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021
Subjects: Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[135] arXiv:2102.05151 (cross-list from cs.SD) [pdf, other]
Title: Enhancing Audio Augmentation Methods with Consistency Learning
Turab Iqbal, Karim Helwani, Arvindh Krishnaswamy, Wenwu Wang
Comments: Accepted to 46th International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2021)
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[136] arXiv:2102.05225 (cross-list from cs.SD) [pdf, other]
Title: Exploring Automatic COVID-19 Diagnosis via voice and symptoms from Crowdsourced Data
Jing Han, Chloë Brown, Jagmohan Chauhan, Andreas Grammenos, Apinan Hasthanasombat, Dimitris Spathis, Tong Xia, Pietro Cicuta, Cecilia Mascolo
Comments: 5 pages, 3 figures, 2 tables, Accepted for publication at ICASSP 2021
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[137] arXiv:2102.05630 (cross-list from cs.SD) [pdf, other]
Title: Voice Cloning: a Multi-Speaker Text-to-Speech Synthesis Approach based on Transfer Learning
Giuseppe Ruggiero, Enrico Zovato, Luigi Di Caro, Vincent Pollet
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[138] arXiv:2102.05749 (cross-list from cs.SD) [pdf, other]
Title: Self-Supervised VQ-VAE for One-Shot Music Style Transfer
Ondřej Cífka, Alexey Ozerov, Umut Şimşekli, Gaël Richard
Comments: ICASSP 2021. Website: this https URL
Journal-ref: ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (2021) 96-100
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
[139] arXiv:2102.05872 (cross-list from cs.SD) [pdf, other]
Title: Onoma-to-wave: Environmental sound synthesis from onomatopoeic words
Yuki Okamoto, Keisuke Imoto, Shinnosuke Takamichi, Ryosuke Yamanishi, Takahiro Fukumori, Yoichi Yamashita
Comments: Accepted to APSIPA Transactions on Signal and Information Processing
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[140] arXiv:2102.05894 (cross-list from cs.SD) [pdf, other]
Title: CASA-Based Speaker Identification Using Cascaded GMM-CNN Classifier in Noisy and Emotional Talking Conditions
Ali Bou Nassif, Ismail Shahin, Shibani Hamsa, Nawel Nemmour, Keikichi Hirose
Comments: Published in Applied Soft Computing journal
Journal-ref: Applied Soft Computing, Elsevier, 2021
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[141] arXiv:2102.06003 (cross-list from cs.SD) [pdf, other]
Title: Language Independent Emotion Quantification using Non linear Modelling of Speech
Uddalok Sarkar, Sayan Nag, Chirayata Bhattacharya, Shankha Sanyal, Archi Banerjee, Ranjan Sengupta, Dipak Ghosh
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[142] arXiv:2102.06034 (cross-list from cs.SD) [pdf, other]
Title: Speech enhancement with mixture-of-deep-experts with clean clustering pre-training
Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot
Comments: arXiv admin note: text overlap with arXiv:1703.09302
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[143] arXiv:2102.06038 (cross-list from cs.SD) [pdf, other]
Title: A Fractal Approach to Characterize Emotions in Audio and Visual Domain: A Study on Cross-Modal Interaction
Sayan Nag, Uddalok Sarkar, Shankha Sanyal, Archi Banerjee, Souparno Roy, Samir Karmakar, Ranjan Sengupta, Dipak Ghosh
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[144] arXiv:2102.06142 (cross-list from cs.SD) [pdf, other]
Title: Multichannel-based learning for audio object extraction
Daniel Arteaga, Jordi Pons
Comments: In proceedings of ICASSP2021. Appendix added
Journal-ref: ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 206-210
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[145] arXiv:2102.06269 (cross-list from eess.IV) [pdf, other]
Title: Disentanglement for audio-visual emotion recognition using multitask setup
Raghuveer Peri, Srinivas Parthasarathy, Charles Bradshaw, Shiva Sundaram
Comments: Accepted for ICASSP 2021, 5 pages
Subjects: Image and Video Processing (eess.IV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[146] arXiv:2102.06283 (cross-list from cs.CL) [pdf, other]
Title: Speech-language Pre-training for End-to-end Spoken Language Understanding
Yao Qian, Ximo Bian, Yu Shi, Naoyuki Kanda, Leo Shen, Zhen Xiao, Michael Zeng
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[147] arXiv:2102.06291 (cross-list from cs.SD) [pdf, other]
Title: A Multi-View Approach To Audio-Visual Speaker Verification
Leda Sarı, Kritika Singh, Jiatong Zhou, Lorenzo Torresani, Nayan Singhal, Yatharth Saraf
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[148] arXiv:2102.06357 (cross-list from cs.SD) [pdf, other]
Title: Contrastive Unsupervised Learning for Speech Emotion Recognition
Mao Li, Bo Yang, Joshua Levy, Andreas Stolcke, Viktor Rozgic, Spyros Matsoukas, Constantinos Papayiannis, Daniel Bone, Chao Wang
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[149] arXiv:2102.06380 (cross-list from cs.CL) [pdf, other]
Title: Neural Inverse Text Normalization
Monica Sunkara, Chaitanya Shivade, Sravan Bodapati, Katrin Kirchhoff
Comments: 5 pages, accepted to ICASSP 2021
Subjects: Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[150] arXiv:2102.06393 (cross-list from eess.SP) [pdf, other]
Title: Mind the beat: detecting audio onsets from EEG recordings of music listening
Ashvala Vinay, Alexander Lerch, Grace Leslie
Comments: to be published in ICASSP 2021 4 figures, 5 pages (4 pages of content + 1 page of references)
Subjects: Signal Processing (eess.SP); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[151] arXiv:2102.06431 (cross-list from cs.SD) [pdf, other]
Title: VARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention
Peng Liu, Yuewen Cao, Songxiang Liu, Na Hu, Guangzhi Li, Chao Weng, Dan Su
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[152] arXiv:2102.06455 (cross-list from cs.SD) [pdf, other]
Title: Deep Sound Field Reconstruction in Real Rooms: Introducing the ISOBEL Sound Field Dataset
Miklas Strøm Kristoffersen, Martin Bo Møller, Pablo Martínez-Nuevo, Jan Østergaard
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[153] arXiv:2102.06467 (cross-list from cs.SD) [pdf, other]
Title: Content-Aware Speaker Embeddings for Speaker Diarisation
G. Sun, D. Liu, C. Zhang, P. C. Woodland
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[154] arXiv:2102.06657 (cross-list from cs.CV) [pdf, other]
Title: End-to-end Audio-visual Speech Recognition with Conformers
Pingchuan Ma, Stavros Petridis, Maja Pantic
Comments: Accepted to ICASSP 2021
Subjects: Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[155] arXiv:2102.06750 (cross-list from cs.CL) [pdf, other]
Title: Do as I mean, not as I say: Sequence Loss Training for Spoken Language Understanding
Milind Rao, Pranav Dheram, Gautam Tiwari, Anirudh Raju, Jasha Droppo, Ariya Rastrow, Andreas Stolcke
Comments: Proc. IEEE ICASSP 2021
Subjects: Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[156] arXiv:2102.06930 (cross-list from cs.SD) [pdf, other]
Title: Deep Convolutional and Recurrent Networks for Polyphonic Instrument Classification from Monophonic Raw Audio Waveforms
Kleanthis Avramidis, Agelos Kratimenos, Christos Garoufis, Athanasia Zlatintsi, Petros Maragos
Comments: 5 pages, 4 figures, 6 tables, to be published in the Proc. of the 46th International Conference on Acoustics, Speech and Signal Processing (ICASSP 2021) @ Toronto, Ontario, Canada
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[157] arXiv:2102.06934 (cross-list from cs.SD) [pdf, other]
Title: Multi-Channel Speech Enhancement using Graph Neural Networks
Panagiotis Tzirakis, Anurag Kumar, Jacob Donley
Journal-ref: Proc. ICASSP 2021
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[158] arXiv:2102.07133 (cross-list from cs.SD) [pdf, other]
Title: Parametric Optimization of Violin Top Plates using Machine Learning
Davide Salvi, Sebastian Gonzalez, Fabio Antonacci, Augusto Sarti
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[159] arXiv:2102.07259 (cross-list from cs.SD) [pdf, other]
Title: Thank you for Attention: A survey on Attention-based Artificial Neural Networks for Automatic Speech Recognition
Priyabrata Karmakar, Shyh Wei Teng, Guojun Lu
Comments: Submitted to IEEE/ACM Trans. on Audio, Speech, and Language Processing
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[160] arXiv:2102.07307 (cross-list from cs.SD) [pdf, other]
Title: I-vector Based Within Speaker Voice Quality Identification on connected speech
Chuyao Feng, Eva van Leer, Mackenzie Lee Curtis, David V. Anderson
Comments: s
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[161] arXiv:2102.07594 (cross-list from cs.CL) [pdf, other]
Title: Fast End-to-End Speech Recognition via Non-Autoregressive Models and Cross-Modal Knowledge Transferring from BERT
Ye Bai, Jiangyan Yi, Jianhua Tao, Zhengkun Tian, Zhengqi Wen, Shuai Zhang
Comments: 14 pages, 7 figures
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[162] arXiv:2102.07896 (cross-list from eess.SP) [pdf, other]
Title: A multispeaker dataset of raw and reconstructed speech production real-time MRI video and 3D volumetric images
Yongwan Lim, Asterios Toutios, Yannick Bliesener, Ye Tian, Sajan Goud Lingala, Colin Vaz, Tanner Sorensen, Miran Oh, Sarah Harper, Weiyi Chen, Yoonjeong Lee, Johannes Töger, Mairym Lloréns Montesserin, Caitlin Smith, Bianca Godinez, Louis Goldstein, Dani Byrd, Krishna S. Nayak, Shrikanth S. Narayanan
Comments: 27 pages, 6 figures, 5 tables, submitted to Nature Scientific Data
Subjects: Signal Processing (eess.SP); Sound (cs.SD); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[163] arXiv:2102.07982 (cross-list from cs.SD) [pdf, other]
Title: Voice Gender Scoring and Independent Acoustic Characterization of Perceived Masculinity and Femininity
Fuling Chen, Roberto Togneri, Murray Maybery, Diana Tan
Comments: 24 pages, 7 figures, journal
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[164] arXiv:2102.07990 (cross-list from eess.SP) [pdf, other]
Title: Through-the-Wall Radar under Electromagnetic Complex Wall: A Deep Learning Approach
Fardin Ghorbani, Hossein Soleimani
Subjects: Signal Processing (eess.SP); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[165] arXiv:2102.08015 (cross-list from cs.SD) [pdf, other]
Title: Improving speech recognition models with small samples for air traffic control systems
Yi Lin, Qin Li, Bo Yang, Zhen Yan, Huachun Tan, Zhengmao Chen
Comments: This work has been accepted by Neurocomputing for publication
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[166] arXiv:2102.08074 (cross-list from cs.SD) [pdf, other]
Title: Semi Supervised Learning For Few-shot Audio Classification By Episodic Triplet Mining
Swapnil Bhosale, Rupayan Chakraborty, Sunil Kumar Kopparapu
Comments: 5 pages
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[167] arXiv:2102.08183 (cross-list from cs.SD) [pdf, other]
Title: Comparison of semi-supervised deep learning algorithms for audio classification
Léo Cances, Etienne Labbé, Thomas Pellegrini
Comments: 9 pages, 5 figures, 5 tables. This is the version 3 of the paper. Contains minor fixes compared to the EURASIP one (which is the version 2 of the paper)
Journal-ref: EURASIP Journal on Audio, Speech, and Music Processing, 2022
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[168] arXiv:2102.08359 (cross-list from cs.SD) [pdf, other]
Title: End-2-End COVID-19 Detection from Breath & Cough Audio
Harry Coppock, Alexander Gaskell, Panagiotis Tzirakis, Alice Baird, Lyn Jones, Björn W. Schuller
Comments: 5 pages
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[169] arXiv:2102.08535 (cross-list from cs.CL) [pdf, other]
Title: ATCSpeechNet: A multilingual end-to-end speech recognition framework for air traffic control systems
Yi Lin, Bo Yang, Linchao Li, Dongyue Guo, Jianwei Zhang, Hu Chen, Yi Zhang
Comments: An improved work based on our previous Interspeech 2020 paper (this https URL)
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[170] arXiv:2102.08551 (cross-list from cs.SD) [pdf, other]
Title: Weighted Recursive Least Square Filter and Neural Network based Residual Echo Suppression for the AEC-Challenge
Ziteng Wang, Yueyue Na, Zhang Liu, Biao Tian, Qiang Fu
Comments: 5 pages, 2 figures, accepted by ICASSP 2021
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[171] arXiv:2102.08575 (cross-list from cs.SD) [pdf, other]
Title: End-to-end lyrics Recognition with Voice to Singing Style Transfer
Sakya Basak, Shrutina Agarwal, Sriram Ganapathy, Naoya Takahashi
Comments: accepted at ICASSP 2021
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[172] arXiv:2102.08833 (cross-list from cs.SD) [pdf, other]
Title: DESED-FL and URBAN-FL: Federated Learning Datasets for Sound Event Detection
David S. Johnson, Wolfgang Lorenz, Michael Taenzer, Stylianos Mimilakis, Sascha Grollmisch, Jakob Abeßer, Hanna Lukashevich
Comments: To be published in EUSIPCO 2021
Subjects: Sound (cs.SD); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[173] arXiv:2102.09202 (cross-list from cs.SD) [pdf, other]
Title: Low Resource Audio-to-Lyrics Alignment From Polyphonic Music Recordings
Emir Demirel, Sven Ahlbäck, Simon Dixon
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[174] arXiv:2102.09281 (cross-list from cs.LG) [pdf, other]
Title: DINO: A Conditional Energy-Based GAN for Domain Translation
Konstantinos Vougioukas, Stavros Petridis, Maja Pantic
Comments: Accepted to ICLR 2021
Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[175] arXiv:2102.09607 (cross-list from cs.LG) [pdf, other]
Title: Modelling Paralinguistic Properties in Conversational Speech to Detect Bipolar Disorder and Borderline Personality Disorder
Bo Wang, Yue Wu, Nemanja Vaci, Maria Liakata, Terry Lyons, Kate E A Saunders
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Total of 208 entries : 1-50 51-100 101-150 126-175 151-200 201-208
Showing up to 50 entries per page: fewer | more | all
  • About
  • Help
  • contact arXivClick here to contact arXiv Contact
  • subscribe to arXiv mailingsClick here to subscribe Subscribe
  • Copyright
  • Privacy Policy
  • Web Accessibility Assistance
  • arXiv Operational Status
    Get status notifications via email or slack