Skip to main content
Cornell University
We gratefully acknowledge support from the Simons Foundation, member institutions, and all contributors. Donate
arxiv logo > cs.SD

Help | Advanced Search

arXiv logo
Cornell University Logo

quick links

  • Login
  • Help Pages
  • About

Sound

Authors and titles for September 2021

Total of 163 entries
Showing up to 2000 entries per page: fewer | more | all
[1] arXiv:2109.00103 [pdf, other]
Title: Automatic non-invasive Cough Detection based on Accelerometer and Audio Signals
Madhurananda Pahar, Igor Miranda, Andreas Diacon, Thomas Niesler
Comments: arXiv admin note: text overlap with arXiv:2102.04997
Journal-ref: Journal of Signal Processing Systems, 2022
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[2] arXiv:2109.00181 [pdf, other]
Title: CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations
Hang Li, Yu Kang, Tianqiao Liu, Wenbiao Ding, Zitao Liu
Comments: The 2021 Conference on Empirical Methods in Natural Language Processing
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI)
[3] arXiv:2109.00237 [pdf, other]
Title: Prior Distribution Design for Music Bleeding-Sound Reduction Based on Nonnegative Matrix Factorization
Yusaku Mizobuchi, Daichi Kitamura, Tomohiko Nakamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo
Comments: Accepted and will be presented at APSIPA2021
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[4] arXiv:2109.00260 [pdf, other]
Title: A Separable Temporal Convolution Neural Network with Attention for Small-Footprint Keyword Spotting
Shenghua Hu, Jing Wang, Yujun Wang, Lidong Yang, Wenjing Yang
Comments: arXiv admin note: text overlap with arXiv:2108.12146
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[5] arXiv:2109.00265 [pdf, other]
Title: Embedding and Beamforming: All-neural Causal Beamformer for Multichannel Speech Enhancement
Andong Li, Wenzhe Liu, Chengshi Zheng, Xiaodong Li
Comments: Submitted to ICASSP 2022, first version
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[6] arXiv:2109.00630 [pdf, other]
Title: A Novel Multi-Centroid Template Matching Algorithm and Its Application to Cough Detection
Shibo Zhang, Ebrahim Nemati, Tousif Ahmed, Md Mahbubur Rahman, Jilong Kuang, Alex Gao
Subjects: Sound (cs.SD); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[7] arXiv:2109.00663 [pdf, other]
Title: Controllable deep melody generation via hierarchical music structure representation
Shuqi Dai, Zeyu Jin, Celso Gomes, Roger B. Dannenberg
Comments: 6 pages, 9 figures, in Proc. of the 22nd Int. Society for Music Information Retrieval Conf.,Online, 2021
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[8] arXiv:2109.00704 [pdf, other]
Title: Multichannel Audio Source Separation with Independent Deeply Learned Matrix Analysis Using Product of Source Models
Takuya Hasumi, Tomohiko Nakamura, Norihiro Takamune, Hiroshi Saruwatari, Daichi Kitamura, Yu Takahashi, Kazunobu Kondo
Comments: 8 pages, 5 figures, accepted for Asia-Pacific Signal and Information Processing Association Annual Summit and Conference 2021 (APSIPA ASC 2021)
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[9] arXiv:2109.00748 [pdf, other]
Title: Binaural Audio Generation via Multi-task Learning
Sijia Li, Shiguang Liu, Dinesh Manocha
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[10] arXiv:2109.01948 [pdf, other]
Title: Network Modulation Synthesis: New Algorithms for Generating Musical Audio Using Autoencoder Networks
Jeremy Hyrkas
Comments: accepted to the International Computer Music Conference 2021 (2020 Selected Papers)
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[11] arXiv:2109.01989 [pdf, other]
Title: The SpeakIn System for VoxCeleb Speaker Recognition Challange 2021
Miao Zhao, Yufeng Ma, Min Liu, Minqiang Xu
Comments: Submitted to INTERSPEECH2021 VoxSRC2021 Workshop
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[12] arXiv:2109.02011 [pdf, other]
Title: A Two-stage Complex Network using Cycle-consistent Generative Adversarial Networks for Speech Enhancement
Guochen Yu, Yutian Wang, Hui Wang, Qin Zhang, Chengshi Zheng
Comments: Accepted by Speech Communication
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[13] arXiv:2109.02047 [pdf, other]
Title: The ByteDance Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2021
Keke Wang, Xudong Mao, Hao Wu, Chen Ding, Chuxiang Shang, Rui Xia, Yuxuan Wang
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[14] arXiv:2109.02051 [pdf, other]
Title: Efficient Attention Branch Network with Combined Loss Function for Automatic Speaker Verification Spoof Detection
Amir Mohammad Rostami, Mohammad Mehdi Homayounpour, Ahmad Nickabadi
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Audio and Speech Processing (eess.AS)
[15] arXiv:2109.02052 [pdf, other]
Title: The Phonexia VoxCeleb Speaker Recognition Challenge 2021 System Description
Josef Slavíček, Albert Swart, Michal Klčo, Niko Brümmer
Comments: Second place in the self-supervised track of VoxSRC-21: VoxCeleb Speaker Recognition Challenge
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[16] arXiv:2109.02096 [pdf, other]
Title: Timbre Transfer with Variational Auto Encoding and Cycle-Consistent Adversarial Networks
Russell Sammut Bonnici, Charalampos Saitis, Martin Benning
Comments: 12 pages, 3 main figures, 4 tables
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[17] arXiv:2109.02472 [pdf, other]
Title: Audio-based Musical Version Identification: Elements and Challenges
Furkan Yesiler, Guillaume Doras, Rachel M. Bittner, Christopher J. Tralie, Joan Serrà
Comments: Accepted to be published in IEEE Signal Processing Magazine
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[18] arXiv:2109.02692 [pdf, other]
Title: Machine Learning: Challenges, Limitations, and Compatibility for Audio Restoration Processes
Owen Casey, Rushit Dave, Naeem Seliya, Evelyn R Sowells Boone
Comments: 6 pages, 2 figures
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[19] arXiv:2109.02763 [pdf, other]
Title: Binaural SoundNet: Predicting Semantics, Depth and Motion with Binaural Sounds
Dengxin Dai, Arun Balajee Vasudevan, Jiri Matas, Luc Van Gool
Comments: Accepted by TPAMI. arXiv admin note: substantial text overlap with arXiv:2003.04210
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[20] arXiv:2109.02773 [pdf, other]
Title: Complementing Handcrafted Features with Raw Waveform Using a Light-weight Auxiliary Model
Zhongwei Teng, Quchen Fu, Jules White, Maria Powell, Douglas C. Schmidt
Subjects: Sound (cs.SD); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[21] arXiv:2109.02774 [pdf, other]
Title: FastAudio: A Learnable Audio Front-End for Spoof Speech Detection
Quchen Fu, Zhongwei Teng, Jules White, Maria Powell, Douglas C. Schmidt
Subjects: Sound (cs.SD); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[22] arXiv:2109.03219 [pdf, other]
Title: Fruit-CoV: An Efficient Vision-based Framework for Speedy Detection and Diagnosis of SARS-CoV-2 Infections Through Recorded Cough Sounds
Long H. Nguyen, Nhat Truong Pham, Van Huong Do, Liu Tai Nguyen, Thanh Tin Nguyen, Van Dung Do, Hai Nguyen, Ngoc Duy Nguyen
Comments: 4 pages
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE); Audio and Speech Processing (eess.AS)
[23] arXiv:2109.03465 [pdf, other]
Title: A Survey of Sound Source Localization with Deep Learning Methods
Pierre-Amaury Grumiaux, Srđan Kitić, Laurent Girin, Alexandre Guérin
Comments: Accepted for publication in The Journal of the Acoustical Society of America
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[24] arXiv:2109.03551 [pdf, other]
Title: Time Alignment using Lip Images for Frame-based Electrolaryngeal Voice Conversion
Yi-Syuan Liou, Wen-Chin Huang, Ming-Chi Yen, Shu-Wei Tsai, Yu-Huai Peng, Tomoki Toda, Yu Tsao, Hsin-Min Wang
Comments: Accepted to APSIPA ASC 2021
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[25] arXiv:2109.03568 [pdf, other]
Title: Beijing ZKJ-NPU Speaker Verification System for VoxCeleb Speaker Recognition Challenge 2021
Li Zhang, Huan Zhao, Qinling Meng, Yanli Chen, Min Liu, Lei Xie
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[26] arXiv:2109.04049 [pdf, other]
Title: BeamTransformer: Microphone Array-based Overlapping Speech Detection
Siqi Zheng, Shiliang Zhang, Weilong Huang, Qian Chen, Hongbin Suo, Ming Lei, Jinwei Feng, Zhijie Yan
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[27] arXiv:2109.04081 [pdf, other]
Title: DeepEMO: Deep Learning for Speech Emotion Recognition
Enkhtogtokh Togootogtokh, Christian Klasen
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[28] arXiv:2109.04658 [pdf, other]
Title: Speech Enhancement by Noise Self-Supervised Rank-Constrained Spatial Covariance Matrix Estimation via Independent Deeply Learned Matrix Analysis
Sota Misawa, Norihiro Takamune, Tomohiko Nakamura, Daichi Kitamura, Hiroshi Saruwatari, Masakazu Une, Shoji Makino
Comments: accepted for APSIPA2021
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[29] arXiv:2109.04783 [pdf, other]
Title: Self-Attention Channel Combinator Frontend for End-to-End Multichannel Far-field Speech Recognition
Rong Gong, Carl Quillen, Dushyant Sharma, Andrew Goderre, José Laínez, Ljubomir Milanović
Comments: In Proceedings of Interspeech 2021
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[30] arXiv:2109.05418 [pdf, other]
Title: Decoupling Magnitude and Phase Estimation with Deep ResUNet for Music Source Separation
Qiuqiang Kong, Yin Cao, Haohe Liu, Keunwoo Choi, Yuxuan Wang
Comments: 6 pages
Journal-ref: International Society for Music Information Retrieval (ISMIR) 2021
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[31] arXiv:2109.05426 [pdf, other]
Title: Zero-Shot Text-to-Speech for Text-Based Insertion in Audio Narration
Chuanxin Tang, Chong Luo, Zhiyuan Zhao, Dacheng Yin, Yucheng Zhao, Wenjun Zeng
Comments: Published in Interspeech'21
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[32] arXiv:2109.06441 [pdf, other]
Title: Structure-Enhanced Pop Music Generation via Harmony-Aware Learning
Xueyao Zhang, Jinchao Zhang, Yao Qiu, Li Wang, Jie Zhou
Comments: Accepted by ACM MM 2022
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[33] arXiv:2109.06459 [pdf, other]
Title: A Machine-learning Framework for Acoustic Design Assessment in Early Design Stages
Reyhane Abarghooie, Zahra Sadat Zomorodian, Mohammad Tahsildoost, Zohreh Shaghaghian
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[34] arXiv:2109.06733 [pdf, other]
Title: Cross-speaker emotion disentangling and transfer for end-to-end speech synthesis
Tao Li, Xinsheng Wang, Qicong Xie, Zhichao Wang, Lei Xie
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[35] arXiv:2109.07623 [pdf, other]
Title: BacHMMachine: An Interpretable and Scalable Model for Algorithmic Harmonization for Four-part Baroque Chorales
Yunyao Zhu, Stephen Hahn, Simon Mak, Yue Jiang, Cynthia Rudin
Comments: 7 pages, 7 figures
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
[36] arXiv:2109.08704 [pdf, other]
Title: Speaker Placement Agnosticism: Improving the Distance-based Amplitude Panning Algorithm
Jacob Sundstrom
Comments: I3DA 2021 International Conference
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[37] arXiv:2109.08839 [pdf, other]
Title: SpeechNAS: Towards Better Trade-off between Latency and Accuracy for Large-Scale Speaker Verification
Wentao Zhu, Tianlong Kong, Shun Lu, Jixiang Li, Dawei Zhang, Feng Deng, Xiaorui Wang, Sen Yang, Ji Liu
Comments: 8 pages, 3 figures, 3 tables. Accepted by ASRU2021
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[38] arXiv:2109.08910 [pdf, other]
Title: MS-SincResNet: Joint learning of 1D and 2D kernels using multi-scale SincNet and ResNet for music genre classification
Pei-Chun Chang, Yong-Sheng Chen, Chang-Hsing Lee
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[39] arXiv:2109.09026 [pdf, other]
Title: Hybrid Data Augmentation and Deep Attention-based Dilated Convolutional-Recurrent Neural Networks for Speech Emotion Recognition
Nhat Truong Pham, Duc Ngoc Minh Dang, Sy Dzung Nguyen
Comments: 12 pages, 16 figures, 6 tables
Subjects: Sound (cs.SD); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[40] arXiv:2109.09227 [pdf, other]
Title: ARCA23K: An audio dataset for investigating open-set label noise
Turab Iqbal, Yin Cao, Andrew Bailey, Mark D. Plumbley, Wenwu Wang
Comments: Accepted to the Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021)
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[41] arXiv:2109.09617 [pdf, other]
Title: TeleMelody: Lyric-to-Melody Generation with a Template-Based Two-Stage Method
Zeqian Ju, Peiling Lu, Xu Tan, Rui Wang, Chen Zhang, Songruoyao Wu, Kejun Zhang, Xiangyang Li, Tao Qin, Tie-Yan Liu
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[42] arXiv:2109.09906 [pdf, other]
Title: Audio Interval Retrieval using Convolutional Neural Networks
Ievgeniia Kuzminykh, Dan Shevchuk, Stavros Shiaeles, Bogdan Ghita
Comments: 20th International Conference on Next Generation Teletraffic and Wired/Wireless Advanced Networks and Systems, NEW2AN 2020 and 13th Conference on the Internet of Things and Smart Spaces, ruSMART 2020
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[43] arXiv:2109.10455 [pdf, other]
Title: An Audio Synthesis Framework Derived from Industrial Process Control
Ashwin Pillay
Comments: 10 pages, 24 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[44] arXiv:2109.10561 [pdf, other]
Title: A Few-Shot Learning Approach for Sound Source Distance Estimation Using Relation Networks
Amirreza Sobhdel, Roozbeh Razavi-Far, Vasile Palade
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[45] arXiv:2109.10608 [pdf, other]
Title: Noisy-to-Noisy Voice Conversion Framework with Denoising Model
Chao Xie, Yi-Chiao Wu, Patrick Lumban Tobing, Wen-Chin Huang, Tomoki Toda
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[46] arXiv:2109.10724 [pdf, other]
Title: Low-Latency Incremental Text-to-Speech Synthesis with Distilled Context Prediction Network
Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari
Comments: Accepted for ASRU2021
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[47] arXiv:2109.11086 [pdf, other]
Title: Scenario Aware Speech Recognition: Advancements for Apollo Fearless Steps & CHiME-4 Corpora
Szu-Jui Chen, Wei Xia, John H.L. Hansen
Comments: Accepted for ASRU 2021
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[48] arXiv:2109.11115 [pdf, other]
Title: Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning
Rui Li, Dong Pu, Minnie Huang, Bill Huang
Comments: 6 pages, 5 figures, Accepted to IEEE ICASSP 2022
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[49] arXiv:2109.11140 [pdf, other]
Title: Joint speaker diarisation and tracking in switching state-space model
Jeremy H. M. Wong, Yifan Gong
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
[50] arXiv:2109.11313 [pdf, other]
Title: Physics-informed neural networks for one-dimensional sound field predictions with parameterized sources and impedance boundaries
Nikolas Borrel-Jensen, Allan P. Engsig-Karup, Cheol-Ho Jeong
Comments: 11 pages, 5 figures, 3 tables
Journal-ref: Jasa Express Letters 2021, Volume 1, Issue 12, pp. 122402
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Computational Physics (physics.comp-ph)
[51] arXiv:2109.11594 [pdf, other]
Title: Implementation of interactive tools for investigating fundamental frequency response of voiced sounds to auditory stimulation
Hideki Kawahara, Toshie Matsui Kohei, Yatabe Ken-Ichi Sakakibara Minoru Tsuzaki Masanori Morise Toshio Irino
Comments: Accepted for APSIPA ASC 2021
Subjects: Sound (cs.SD); Human-Computer Interaction (cs.HC); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
[52] arXiv:2109.11782 [pdf, other]
Title: Causal Analysis of Carnatic Music: A Preliminary Study
Abhsihek Nandekar, Preeth Khona, Rajani M. B., Anindya Sinha, Nithin Nagaraj
Comments: 22 pages, 12 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[53] arXiv:2109.11946 [pdf, other]
Title: Evaluating X-vector-based Speaker Anonymization under White-box Assessment
Pierre Champion (Inria), Denis Jouvet (Inria), Anthony Larcher (LIUM)
Journal-ref: 23rd International Conference on Speech and Computer - SPECOM 2021, Sep 2021, Saint Petersburg, Russia
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Audio and Speech Processing (eess.AS)
[54] arXiv:2109.12014 [pdf, other]
Title: A data acquisition setup for data driven acoustic design
Romana Rust, Achilleas Xydis, Kurt Heutschi, Nathanaël Perraudin, Gonzalo Casas, Chaoyu Du, Jürgen Strauss, Kurt Eggenschwiler, Fernando Perez-Cruz, Fabio Gramazio, Matthias Kohler
Journal-ref: Building Acoustics. February 2021
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[55] arXiv:2109.12056 [pdf, other]
Title: Parameterized Channel Normalization for Far-field Deep Speaker Verification
Xuechen Liu, Md Sahidullah, Tomi Kinnunen
Comments: Accepted for publication at ASRU 2021
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[56] arXiv:2109.12058 [pdf, other]
Title: Optimized Power Normalized Cepstral Coefficients towards Robust Deep Speaker Verification
Xuechen Liu, Md Sahidullah, Tomi Kinnunen
Comments: Accepted for publication at ASRU 2021
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[57] arXiv:2109.12471 [pdf, other]
Title: Rendering Spatial Sound for Interoperable Experiences in the Audio Metaverse
Jean-Marc Jot, Rémi Audfray, Mark Hertensteiner, Brian Schmidt
Comments: International Conference on Immersive and 3D Audio (i3DA), September 2021
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[58] arXiv:2109.12475 [pdf, other]
Title: General Theory of Music by Icosahedron 3: Musical invariant and Melakarta raga
Yusuke Imai
Comments: 31 pages, 34 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[59] arXiv:2109.12591 [pdf, other]
Title: Joint magnitude estimation and phase recovery using Cycle-in-Cycle GAN for non-parallel speech enhancement
Guochen Yu, Andong Li, Yutian Wang, Yinuo Guo, Hui Wang, Chengshi Zheng
Comments: Accecpted by ICASSP 2022
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[60] arXiv:2109.12690 [pdf, other]
Title: Soundata: A Python library for reproducible use of audio datasets
Magdalena Fuentes, Justin Salamon, Pablo Zinemanas, Martín Rocamora, Genís Paja, Irán R. Román, Marius Miron, Xavier Serra, Juan Pablo Bello
Subjects: Sound (cs.SD); Databases (cs.DB); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[61] arXiv:2109.13072 [pdf, other]
Title: Estimating Angle of Arrival (AoA) of multiple Echoes in a Steering Vector Space
Yu-Lin Wei, Romit Roy Choudhury
Comments: 14 pages, 20 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[62] arXiv:2109.13094 [pdf, other]
Title: Inferring Facing Direction from Voice Signals
Yu-Lin Wei, Rui Li, Abhinav Mehrotra, Romit Roy Choudhury, Nic Lane
Comments: 12 pages, 16 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[63] arXiv:2109.13496 [pdf, other]
Title: FastMVAE2: On improving and accelerating the fast variational autoencoder-based source separation algorithm for determined mixtures
Li Li, Hirokazu Kameoka, Shoji Makino
Comments: submit to IEEE/ACM TASLP, under review
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[64] arXiv:2109.13675 [pdf, other]
Title: FlowVocoder: A small Footprint Neural Vocoder based Normalizing flow for Speech Synthesis
Manh Luong, Viet Anh Tran
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[65] arXiv:2109.13731 [pdf, other]
Title: VoiceFixer: Toward General Speech Restoration with Neural Vocoder
Haohe Liu, Qiuqiang Kong, Qiao Tian, Yan Zhao, DeLiang Wang, Chuanzeng Huang, Yuxuan Wang
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[66] arXiv:2109.13821 [pdf, other]
Title: Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, Mikhail Kudinov, Jiansheng Wei
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Machine Learning (stat.ML)
[67] arXiv:2109.14508 [pdf, other]
Title: Cross-domain Semi-Supervised Audio Event Classification Using Contrastive Regularization
Donmoon Lee, Kyogu Lee
Comments: 5 pages, 3 figures, and 2 tables. Accepted paper at IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2021
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[68] arXiv:2109.14705 [pdf, other]
Title: Adaptive Approach For Sparse Representations Using The Locally Competitive Algorithm For Audio
Soufiyan Bahadi, Jean Rouat, Éric Plourde
Comments: To be published at IEEE Machine Learning for Signal Processing 2021
Journal-ref: 2021 IEEE 31st International Workshop on Machine Learning for Signal Processing (MLSP)
Subjects: Sound (cs.SD); Computation and Language (cs.CL)
[69] arXiv:2109.14797 [pdf, other]
Title: Emergency Vehicles Audio Detection and Localization in Autonomous Driving
Hongyi Sun, Xinyi Liu, Kecheng Xu, Jinghao Miao, Qi Luo
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Robotics (cs.RO); Audio and Speech Processing (eess.AS)
[70] arXiv:2109.15053 [pdf, other]
Title: Fine-tuning wav2vec2 for speaker recognition
Nik Vaessen, David A. van Leeuwen
Comments: accepted to ICASSP 2022
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[71] arXiv:2109.15188 [pdf, other]
Title: Assessing Algorithmic Biases for Musical Version Identification
Furkan Yesiler, Marius Miron, Joan Serrà, Emilia Gómez
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[72] arXiv:2109.00281 (cross-list from cs.CR) [pdf, other]
Title: Benchmarking and challenges in security and privacy for voice biometrics
Jean-Francois Bonastre, Hector Delgado, Nicholas Evans, Tomi Kinnunen, Kong Aik Lee, Xuechen Liu, Andreas Nautsch, Paul-Gauthier Noe, Jose Patino, Md Sahidullah, Brij Mohan Lal Srivastava, Massimiliano Todisco, Natalia Tomashenko, Emmanuel Vincent, Xin Wang, Junichi Yamagishi
Comments: Submitted to the symposium of the ISCA Security & Privacy in Speech Communications (SPSC) special interest group
Subjects: Cryptography and Security (cs.CR); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[73] arXiv:2109.00393 (cross-list from cs.NE) [pdf, other]
Title: Mean absorption estimation from room impulse responses using virtually supervised learning
Cédric Foy (UMRAE ), Antoine Deleforge (MULTISPEECH), Diego Di Carlo (PANAMA)
Journal-ref: Journal of the Acoustical Society of America, Acoustical Society of America, 2021, 150 (2), pp.1286-1299
Subjects: Neural and Evolutionary Computing (cs.NE); Sound (cs.SD); Audio and Speech Processing (eess.AS); Classical Physics (physics.class-ph)
[74] arXiv:2109.00535 (cross-list from eess.AS) [pdf, other]
Title: ASVspoof 2021: Automatic Speaker Verification Spoofing and Countermeasures Challenge Evaluation Plan
Héctor Delgado, Nicholas Evans, Tomi Kinnunen, Kong Aik Lee, Xuechen Liu, Andreas Nautsch, Jose Patino, Md Sahidullah, Massimiliano Todisco, Xin Wang, Junichi Yamagishi
Comments: this http URL
Subjects: Audio and Speech Processing (eess.AS); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Sound (cs.SD)
[75] arXiv:2109.00537 (cross-list from eess.AS) [pdf, other]
Title: ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection
Junichi Yamagishi, Xin Wang, Massimiliano Todisco, Md Sahidullah, Jose Patino, Andreas Nautsch, Xuechen Liu, Kong Aik Lee, Tomi Kinnunen, Nicholas Evans, Héctor Delgado
Comments: Accepted to the ASVspoof 2021 Workshop
Subjects: Audio and Speech Processing (eess.AS); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Sound (cs.SD)
[76] arXiv:2109.00577 (cross-list from cs.LG) [pdf, other]
Title: FaVoA: Face-Voice Association Favours Ambiguous Speaker Detection
Hugo Carneiro, Cornelius Weber, Stefan Wermter
Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[77] arXiv:2109.00627 (cross-list from cs.CL) [pdf, other]
Title: Tree-constrained Pointer Generator for End-to-end Contextual Speech Recognition
Guangzhi Sun, Chao Zhang, Philip C. Woodland
Comments: To appear in ASRU 2021
Subjects: Computation and Language (cs.CL); Sound (cs.SD)
[78] arXiv:2109.00648 (cross-list from cs.CL) [pdf, other]
Title: The VoicePrivacy 2020 Challenge: Results and findings
Natalia Tomashenko, Xin Wang, Emmanuel Vincent, Jose Patino, Brij Mohan Lal Srivastava, Paul-Gauthier Noé, Andreas Nautsch, Nicholas Evans, Junichi Yamagishi, Benjamin O'Brien, Anaïs Chanclu, Jean-François Bonastre, Massimiliano Todisco, Mohamed Maouche
Comments: Submitted to the Special Issue on Voice Privacy (Computer Speech and Language Journal - Elsevier); under review
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[79] arXiv:2109.00913 (cross-list from eess.AS) [pdf, other]
Title: Physiological-Physical Feature Fusion for Automatic Voice Spoofing Detection
Junxiao Xue, Hao Zhou, Yabo Wang
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Image and Video Processing (eess.IV)
[80] arXiv:2109.00928 (cross-list from eess.AS) [pdf, other]
Title: Speaker-Conditioned Hierarchical Modeling for Automated Speech Scoring
Yaman Kumar Singla, Avykat Gupta, Shaurya Bagga, Changyou Chen, Balaji Krishnamurthy, Rajiv Ratn Shah
Comments: Published in CIKM 2021
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
[81] arXiv:2109.00962 (cross-list from eess.AS) [pdf, other]
Title: You Only Hear Once: A YOLO-like Algorithm for Audio Segmentation and Sound Event Detection
Satvik Venkatesh, David Moffat, Eduardo Reck Miranda
Comments: 19 pages, 4 figures, 8 tables. Added more experimental validation and background information. Published in Applied Sciences
Journal-ref: Appl.Sci. 12 (2022) 3293
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[82] arXiv:2109.01163 (cross-list from eess.AS) [pdf, other]
Title: Efficient conformer: Progressive downsampling and grouped attention for automatic speech recognition
Maxime Burchi, Valentin Vielzeuf
Journal-ref: ASRU 2021, Dec 2021, Cartagena, Colombia
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)
[83] arXiv:2109.01164 (cross-list from eess.AS) [pdf, other]
Title: Scalable Data Annotation Pipeline for High-Quality Large Speech Datasets Development
Mingkuan Liu, Chi Zhang, Hua Xing, Chao Feng, Monchu Chen, Judith Bishop, Grace Ngapo
Comments: Submitted to NeurIPS 2021 Datasets and Benchmarks Track (Round 2)
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[84] arXiv:2109.01568 (cross-list from eess.AS) [pdf, other]
Title: Phone Duration Modeling for Speaker Age Estimation in Children
Prashanth Gurunath Shivakumar, Somer Bishop, Catherine Lord, Shrikanth Narayanan
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[85] arXiv:2109.01607 (cross-list from eess.AS) [pdf, other]
Title: Musical Tempo Estimation Using a Multi-scale Network
Xiaoheng Sun, Qiqi He, Yongwei Gao, Wei Li
Comments: Accepted by ISMIR 2021
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[86] arXiv:2109.01766 (cross-list from cs.CR) [pdf, other]
Title: SEC4SR: A Security Analysis Platform for Speaker Recognition
Guangke Chen, Zhe Zhao, Fu Song, Sen Chen, Lingling Fan, Yang Liu
Subjects: Cryptography and Security (cs.CR); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[87] arXiv:2109.02002 (cross-list from eess.AS) [pdf, other]
Title: The DKU-DukeECE-Lenovo System for the Diarization Task of the 2021 VoxCeleb Speaker Recognition Challenge
Weiqing Wang, Danwei Cai, Qingjian Lin, Lin Yang, Junjie Wang, Jin Wang, Ming Li
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[88] arXiv:2109.02549 (cross-list from eess.AS) [pdf, other]
Title: XMUSPEECH System for VoxCeleb Speaker Recognition Challenge 2021
Jie Wang, Fuchuang Tong, Zhicong Chen, Lin Li, Qingyang Hong, Haodong Zhou
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[89] arXiv:2109.02576 (cross-list from eess.AS) [pdf, other]
Title: Improving Speaker Identification for Shared Devices by Adapting Embeddings to Speaker Subsets
Zhenning Tan, Yuguang Yang, Eunjung Han, Andreas Stolcke
Comments: Submitted to ASRU 2021
Journal-ref: Proc. IEEE Automatic Speech Recognition and Understanding Workshop, Dec. 2021, pp. 1124-1131
Subjects: Audio and Speech Processing (eess.AS); Cryptography and Security (cs.CR); Sound (cs.SD)
[90] arXiv:2109.02853 (cross-list from eess.AS) [pdf, other]
Title: The DKU-DukeECE System for the Self-Supervision Speaker Verification Task of the 2021 VoxCeleb Speaker Recognition Challenge
Danwei Cai, Ming Li
Comments: arXiv admin note: text overlap with arXiv:2010.14751
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[91] arXiv:2109.02993 (cross-list from cs.CV) [pdf, other]
Title: Evaluation of an Audio-Video Multimodal Deepfake Dataset using Unimodal and Multimodal Detectors
Hasam Khalid, Minha Kim, Shahroz Tariq, Simon S. Woo
Comments: 2 Figures, 2 Tables, Accepted for publication at the 1st Workshop on Synthetic Multimedia - Audiovisual Deepfake Generation and Detection (ADGD '21) at ACM MM 2021
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[92] arXiv:2109.03264 (cross-list from cs.CL) [pdf, other]
Title: Text-Free Prosody-Aware Generative Spoken Language Modeling
Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu-Anh Nguyen, Morgane Rivière, Abdelrahman Mohamed, Emmanuel Dupoux, Wei-Ning Hsu
Comments: ACL 2022
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[93] arXiv:2109.03275 (cross-list from eess.AS) [pdf, other]
Title: A New Non-Negative Matrix Co-Factorisation Approach for Noisy Neonatal Chest Sound Separation
Ethan Grooby, Jinyuan He, Davood Fattahi, Lindsay Zhou, Arrabella King, Ashwin Ramanathan, Atul Malhotra, Guy A. Dumont, Faezeh Marzbanrad
Comments: 6 pages, 2 figures. To appear as conference paper at 43rd Annual International Conference of the IEEE Engineering in Medicine and Biology Society, 1st-5th November 2021
Journal-ref: 2021 43rd Annual International Conference of the IEEE Engineering in Medicine Biology Society (EMBC)
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Signal Processing (eess.SP)
[94] arXiv:2109.03277 (cross-list from cs.CL) [pdf, other]
Title: A Dual-Decoder Conformer for Multilingual Speech Recognition
Krishna D N
Comments: 5 pages
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[95] arXiv:2109.03381 (cross-list from cs.CL) [pdf, other]
Title: Self-supervised Contrastive Cross-Modality Representation Learning for Spoken Question Answering
Chenyu You, Nuo Chen, Yuexian Zou
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[96] arXiv:2109.03439 (cross-list from eess.AS) [pdf, other]
Title: Referee: Towards reference-free cross-speaker style transfer with low-quality data for expressive speech synthesis
Songxiang Liu, Shan Yang, Dan Su, Dong Yu
Comments: 7 pages, preprint
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[97] arXiv:2109.03454 (cross-list from cs.LG) [pdf, other]
Title: Signal-domain representation of symbolic music for learning embedding spaces
Mathieu Prang (IRCAM), Philippe Esling
Journal-ref: The 2020 Joint Conference on AI Music Creativity, Oct 2020, Stockholm, Sweden
Subjects: Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[98] arXiv:2109.03969 (cross-list from cs.CL) [pdf, other]
Title: Multilingual Speech Recognition for Low-Resource Indian Languages using Multi-Task conformer
Krishna D N
Comments: 5 pages. Rejected from Interspeech 2021. arXiv admin note: substantial text overlap with arXiv:2109.03277
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[99] arXiv:2109.04070 (cross-list from eess.AS) [pdf, other]
Title: The IDLAB VoxCeleb Speaker Recognition Challenge 2021 System Description
Jenthe Thienpondt, Brecht Desplanques, Kris Demuynck
Comments: arXiv admin note: substantial text overlap with arXiv:2104.02370
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[100] arXiv:2109.04241 (cross-list from eess.AS) [pdf, other]
Title: Robust single- and multi-loudspeaker least-squares-based equalization for hearing devices
Henning Schepker, Florian Denk, Birger Kollmeier, Simon Doclo
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[101] arXiv:2109.04411 (cross-list from eess.AS) [pdf, other]
Title: Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
Hirofumi Inaguma, Yosuke Higuchi, Kevin Duh, Tatsuya Kawahara, Shinji Watanabe
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[102] arXiv:2109.04544 (cross-list from eess.AS) [pdf, other]
Title: Directional MCLP Analysis and Reconstruction for Spatial Speech Communication
Srikanth Raj Chetupalli, Thippur V. Sreenivas
Comments: The manuscript is submitted as a full paper to IEEE/ACM Transactions on Audio, Speech and Language Processing
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Signal Processing (eess.SP)
[103] arXiv:2109.04894 (cross-list from eess.AS) [pdf, other]
Title: Large-vocabulary Audio-visual Speech Recognition in Noisy Environments
Wentao Yu, Steffen Zeiler, Dorothea Kolossa
Journal-ref: The IEEE 23rd International Workshop on Multimedia Signal Processing (MMSP), 2021
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Image and Video Processing (eess.IV)
[104] arXiv:2109.05056 (cross-list from cs.CL) [pdf, other]
Title: Speaker Turn Modeling for Dialogue Act Classification
Zihao He, Leili Tavabi, Kristina Lerman, Mohammad Soleymani
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[105] arXiv:2109.05092 (cross-list from eess.AS) [pdf, other]
Title: Remember the context! ASR slot error correction through memorization
Dhanush Bekal, Ashish Shenoy, Monica Sunkara, Sravan Bodapati, Katrin Kirchhoff
Comments: 8 pages, 3 figures, 4 tables, Accepted to ASRU 2021
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[106] arXiv:2109.05172 (cross-list from eess.AS) [pdf, other]
Title: Incorporating Real-world Noisy Speech in Neural-network-based Speech Enhancement Systems
Yangyang Xia, Buye Xu, Anurag Kumar
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[107] arXiv:2109.05494 (cross-list from cs.CL) [pdf, other]
Title: Unsupervised Domain Adaptation Schemes for Building ASR in Low-resource Languages
Anoop C S, Prathosh A P, A G Ramakrishnan
Comments: Submitted to ASRU 2021
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[108] arXiv:2109.05977 (cross-list from eess.AS) [pdf, other]
Title: Studying squeeze-and-excitation used in CNN for speaker verification
Mickael Rouvier, Pierre-Michel Bousquet
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[109] arXiv:2109.06112 (cross-list from cs.CL) [pdf, other]
Title: Beyond Isolated Utterances: Conversational Emotion Recognition
Raghavendra Pappagari, Piotr Żelasko, Jesús Villalba, Laureano Moro-Velazquez, Najim Dehak
Comments: Accepted for ASRU 2021
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[110] arXiv:2109.06171 (cross-list from eess.AS) [pdf, other]
Title: In-filter Computing For Designing Ultra-light Acoustic Pattern Recognizers
Abhishek Ramdas Nair, Shantanu Chakrabartty, Chetan Singh Thakur
Comments: in IEEE Internet of Things Journal
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE); Sound (cs.SD); Systems and Control (eess.SY)
[111] arXiv:2109.06483 (cross-list from eess.AS) [pdf, other]
Title: Overlap-aware low-latency online speaker diarization based on end-to-end local segmentation
Juan M. Coria, Hervé Bredin, Sahar Ghannay, Sophie Rosset
Comments: To appear in ASRU 2021. Code available at this https URL
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[112] arXiv:2109.06684 (cross-list from cs.CL) [pdf, other]
Title: Non-autoregressive Transformer with Unified Bidirectional Decoder for Automatic Speech Recognition
Chuan-Fei Zhang, Yan Liu, Tian-Hao Zhang, Song-Lu Chen, Feng Chen, Xu-Cheng Yin
Comments: 5 pages, 3 figures. Paper submitted to ICASSP 2022
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[113] arXiv:2109.06824 (cross-list from eess.AS) [pdf, other]
Title: Self-Supervised Metric Learning With Graph Clustering For Speaker Diarization
Prachi Singh, Sriram Ganapathy
Comments: 8 pages, Accepted in ASRU 2021
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[114] arXiv:2109.06870 (cross-list from cs.CL) [pdf, other]
Title: Performance-Efficiency Trade-offs in Unsupervised Pre-training for Speech Recognition
Felix Wu, Kwangyoun Kim, Jing Pan, Kyu Han, Kilian Q. Weinberger, Yoav Artzi
Comments: Code available at this https URL
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[115] arXiv:2109.06912 (cross-list from eess.AS) [pdf, other]
Title: fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
Changhan Wang, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Ann Lee, Peng-Jen Chen, Jiatao Gu, Juan Pino
Comments: Accepted to EMNLP 2021 Demo
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[116] arXiv:2109.06952 (cross-list from cs.CL) [pdf, other]
Title: Residual Adapters for Parameter-Efficient ASR Adaptation to Atypical and Accented Speech
Katrin Tomanek, Vicky Zayats, Dirk Padfield, Kara Vaillancourt, Fadi Biadsy
Comments: Accepted to EMNLP 2021
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[117] arXiv:2109.07149 (cross-list from cs.MM) [pdf, other]
Title: Fusion with Hierarchical Graphs for Mulitmodal Emotion Recognition
Shuyun Tang, Zhaojie Luo, Guoshun Nan, Yuichiro Yoshikawa, Ishiguro Hiroshi
Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD)
[118] arXiv:2109.07274 (cross-list from eess.AS) [pdf, other]
Title: Binaural rendering from microphone array signals of arbitrary geometry
Naoto Iijima, Shoichi Koyama, Hiroshi Saruwatari
Comments: The following article has been accepted by Journal of the Acoustical Society of America (JASA). After it is published, it will be found at this http URL
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[119] arXiv:2109.07327 (cross-list from eess.AS) [pdf, other]
Title: Improving Streaming Transformer Based ASR Under a Framework of Self-supervised Learning
Songjun Cao, Yueteng Kang, Yanzhe Fu, Xiaoshuo Xu, Sining Sun, Yike Zhang, Long Ma
Comments: INTERSPEECH2021
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[120] arXiv:2109.07349 (cross-list from eess.AS) [pdf, other]
Title: Improving Accent Identification and Accented Speech Recognition Under a Framework of Self-supervised Learning
Keqi Deng, Songjun Cao, Long Ma
Comments: INTERSPEECH2021
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[121] arXiv:2109.07513 (cross-list from cs.CL) [pdf, other]
Title: Tied & Reduced RNN-T Decoder
Rami Botros (1), Tara N. Sainath (1), Robert David (1), Emmanuel Guzman (1), Wei Li (1), Yanzhang He (1) ((1) Google Inc. USA)
Journal-ref: Proc. Interspeech 2021, 4563-4567
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[122] arXiv:2109.07750 (cross-list from eess.AS) [pdf, other]
Title: Utterance-level neural confidence measure for end-to-end children speech recognition
Wei Liu, Tan Lee
Comments: accepted by ASRU 2021
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[123] arXiv:2109.07846 (cross-list from cs.LG) [pdf, other]
Title: Telehealthcare and Telepathology in Pandemic: A Noninvasive, Low-Cost Micro-Invasive and Multimodal Real-Time Online Application for Early Diagnosis of COVID-19 Infection
Abdullah Bin Shams, Md. Mohsin Sarker Raihan, Md. Mohi Uddin Khan, Ocean Monjur, Rahat Bin Preo
Comments: 32 Pages. This article has been submitted for review to a prestigious journal
Subjects: Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS); Biomolecules (q-bio.BM)
[124] arXiv:2109.07931 (cross-list from eess.AS) [pdf, other]
Title: DDS: A new device-degraded speech dataset for speech enhancement
Haoyu Li, Junichi Yamagishi
Comments: Submitted to Interspeech 2022
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[125] arXiv:2109.07940 (cross-list from eess.AS) [pdf, other]
Title: PDAugment: Data Augmentation by Pitch and Duration Adjustments for Automatic Lyrics Transcription
Chen Zhang, Jiaxing Yu, LuChin Chang, Xu Tan, Jiawei Chen, Tao Qin, Kejun Zhang
Comments: 7 pages
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[126] arXiv:2109.08007 (cross-list from cs.MM) [pdf, other]
Title: Graph Fourier Transform based Audio Zero-watermarking
Longting Xu, Daiyu Huang, Syed Faham Ali Zaidi, Abdul Rauf, Rohan Kumar Das
Subjects: Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[127] arXiv:2109.08125 (cross-list from eess.AS) [pdf, other]
Title: NORESQA: A Framework for Speech Quality Assessment using Non-Matching References
Pranay Manocha, Buye Xu, Anurag Kumar
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[128] arXiv:2109.08555 (cross-list from eess.AS) [pdf, other]
Title: Continuous Streaming Multi-Talker ASR with Dual-path Transducers
Desh Raj, Liang Lu, Zhuo Chen, Yashesh Gaur, Jinyu Li
Comments: Accepted for publication at IEEE ICASSP 2022
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[129] arXiv:2109.08710 (cross-list from eess.AS) [pdf, other]
Title: On-device neural speech synthesis
Sivanand Achanta, Albert Antony, Ladan Golipour, Jiangchuan Li, Tuomo Raitio, Ramya Rasipuram, Francesco Rossi, Jennifer Shi, Jaimin Upadhyay, David Winarsky, Hepeng Zhang
Comments: 7 pages 2 figures, accepted to ASRU 2021
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Performance (cs.PF); Sound (cs.SD)
[130] arXiv:2109.08867 (cross-list from cs.CV) [pdf, other]
Title: V-SlowFast Network for Efficient Visual Sound Separation
Lingyu Zhu, Esa Rahtu
Comments: total 21 pages: main paper 8 pages, references 3 pages, supplementary material 10 pages
Subjects: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[131] arXiv:2109.08870 (cross-list from eess.AS) [pdf, other]
Title: Fast query-by-example speech search using separable model
Yuguang Yang, Yu Pan, Xin Dong, Minqiang Xu
Comments: 8pages, 8 figures
Subjects: Audio and Speech Processing (eess.AS); Information Retrieval (cs.IR); Sound (cs.SD)
[132] arXiv:2109.09598 (cross-list from cs.CR) [pdf, other]
Title: "Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World
Emily Wenger, Max Bronckers, Christian Cianfarani, Jenna Cryan, Angela Sha, Haitao Zheng, Ben Y. Zhao
Comments: 13 pages
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[133] arXiv:2109.09674 (cross-list from eess.AS) [pdf, other]
Title: Improving Text-Independent Speaker Verification with Auxiliary Speakers Using Graph
Jingyu Li, Si-Ioi Ng, Tan Lee
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[134] arXiv:2109.09684 (cross-list from eess.SP) [pdf, other]
Title: Development of In Situ Acoustic Instruments for The Aquatic Environment Study
Aleksandr N. Grekov (1), Nikolay A. Grekov (1), Evgeniy Sychov (1), K.A. Kuzmin (1) ((1) Institute of Natural and Technical Systems)
Comments: 8 pages, 3 figures
Journal-ref: Monitoring systems of environment 2 (2019): 22-29
Subjects: Signal Processing (eess.SP); Sound (cs.SD); Audio and Speech Processing (eess.AS); Atmospheric and Oceanic Physics (physics.ao-ph)
[135] arXiv:2109.10107 (cross-list from cs.CL) [pdf, other]
Title: On the Difficulty of Segmenting Words with Attention
Ramon Sanabria, Hao Tang, Sharon Goldwater
Comments: Accepted at the "Workshop on Insights from Negative Results in NLP" (EMNLP 2021)
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[136] arXiv:2109.10252 (cross-list from cs.LG) [pdf, other]
Title: Audiomer: A Convolutional Transformer For Keyword Spotting
Surya Kant Sahu, Sai Mitheran, Juhi Kamdar, Meet Gandhi
Comments: The results and claims made are incorrect due to data leakage and an erroneous split of datasets
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[137] arXiv:2109.10598 (cross-list from cs.LG) [pdf, other]
Title: Diarisation using location tracking with agglomerative clustering
Jeremy H. M. Wong, Igor Abramovski, Xiong Xiao, Yifan Gong
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[138] arXiv:2109.11010 (cross-list from cs.CL) [pdf, other]
Title: Alzheimers Dementia Detection using Acoustic & Linguistic features and Pre-Trained BERT
Akshay Valsaraj, Ithihas Madala, Nikhil Garg, Veeky Baths
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
[139] arXiv:2109.11641 (cross-list from eess.AS) [pdf, other]
Title: Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection
Wei Xia, Han Lu, Quan Wang, Anshuman Tripathi, Yiling Huang, Ignacio Lopez Moreno, Hasim Sak
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[140] arXiv:2109.11680 (cross-list from cs.CL) [pdf, other]
Title: Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
Qiantong Xu, Alexei Baevski, Michael Auli
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
[141] arXiv:2109.11955 (cross-list from cs.CV) [pdf, other]
Title: Visual Scene Graphs for Audio Source Separation
Moitreya Chatterjee, Jonathan Le Roux, Narendra Ahuja, Anoop Cherian
Comments: Accepted at ICCV 2021
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[142] arXiv:2109.12804 (cross-list from eess.AS) [pdf, other]
Title: Fast-MD: Fast Multi-Decoder End-to-End Speech Translation with Non-Autoregressive Hidden Intermediates
Hirofumi Inaguma, Siddharth Dalmia, Brian Yan, Shinji Watanabe
Comments: Accepted at IEEE ASRU 2021
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[143] arXiv:2109.13217 (cross-list from cs.CL) [pdf, other]
Title: Challenges and Opportunities of Speech Recognition for Bengali Language
M. F. Mridha, Abu Quwsar Ohi, Md. Abdul Hamid, Muhammad Mostafa Monowar
Comments: Accepted in Artificial Intelligence Review
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[144] arXiv:2109.13226 (cross-list from eess.AS) [pdf, other]
Title: BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
Yu Zhang, Daniel S. Park, Wei Han, James Qin, Anmol Gulati, Joel Shor, Aren Jansen, Yuanzhong Xu, Yanping Huang, Shibo Wang, Zongwei Zhou, Bo Li, Min Ma, William Chan, Jiahui Yu, Yongqiang Wang, Liangliang Cao, Khe Chai Sim, Bhuvana Ramabhadran, Tara N. Sainath, Françoise Beaufays, Zhifeng Chen, Quoc V. Le, Chung-Cheng Chiu, Ruoming Pang, Yonghui Wu
Comments: 14 pages, 7 figures, 13 tables; v2: minor corrections, reference baselines and bibliography updated; v3: corrections based on reviewer feedback, bibliography updated
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
[145] arXiv:2109.13354 (cross-list from cs.MM) [pdf, other]
Title: Audio-to-Image Cross-Modal Generation
Maciej Żelaszczyk, Jacek Mańdziuk
Journal-ref: International Joint Conference on Neural Networks, IJCNN 2022, Padua, Italy, 1-8
Subjects: Multimedia (cs.MM); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[146] arXiv:2109.13425 (cross-list from eess.AS) [pdf, other]
Title: The JHU submission to VoxSRC-21: Track 3
Jejin Cho, Jesus Villalba, Najim Dehak
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[147] arXiv:2109.13510 (cross-list from cs.LG) [pdf, other]
Title: VoxCeleb Enrichment for Age and Gender Recognition
Khaled Hechmi, Trung Ngo Trong, Ville Hautamaki, Tomi Kinnunen
Comments: Accepted for presentation at ASRU 2021; repository: this https URL
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[148] arXiv:2109.13673 (cross-list from cs.CL) [pdf, other]
Title: Nana-HDR: A Non-attentive Non-autoregressive Hybrid Model for TTS
Shilun Lin, Wenchao Su, Li Meng, Fenglong Xie, Xinhui Li, Li Lu
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[149] arXiv:2109.13714 (cross-list from eess.AS) [pdf, other]
Title: MSR-NV: Neural Vocoder Using Multiple Sampling Rates
Kentaro Mitsui, Kei Sawada
Comments: 6 pages including supplement, 3 figures, accepted for INTERSPEECH 2022. Audio samples: this https URL
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[150] arXiv:2109.13815 (cross-list from eess.AS) [pdf, other]
Title: Articulatory Coordination for Speech Motor Tracking in Huntington Disease
Matthew Perez, Amrit Romana, Angela Roberts, Noelle Carlozzi, Jennifer Ann Miner, Praveen Dayalu, Emily Mower Provost
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[151] arXiv:2109.14061 (cross-list from eess.AS) [pdf, other]
Title: The impact of non-target events in synthetic soundscapes for sound event detection
Francesca Ronchini, Romain Serizel, Nicolas Turpault, Samuele Cornell
Journal-ref: Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021)
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[152] arXiv:2109.14200 (cross-list from eess.AS) [pdf, other]
Title: Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
Khazar Khorrami, Okko Räsänen
Comments: Final manuscript published in Language Development Research under CC BY-NC-SA 4.0. Pre-print redistributed through arXiv with permission. Replaces corrupted PsyArXiv pre-print repository at this https URL
Journal-ref: Language Development Research, 1(1), 123-191 (2021)
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD)
[153] arXiv:2109.14357 (cross-list from eess.AS) [pdf, other]
Title: Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
Jakob Poncelet, Hugo Van hamme
Comments: To be published in the 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2021)
Journal-ref: 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pp. 169-176
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[154] arXiv:2109.14370 (cross-list from eess.AS) [pdf, other]
Title: Objective-oriented method for uniformation of various directivity representations
Adam Szwajcowski
Comments: Author's Accepted Manuscript from 151st AES Convention
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[155] arXiv:2109.14420 (cross-list from cs.CL) [pdf, other]
Title: FastCorrect 2: Fast Error Correction on Multiple Candidates for Automatic Speech Recognition
Yichong Leng, Xu Tan, Rui Wang, Linchen Zhu, Jin Xu, Wenjie Liu, Linquan Liu, Tao Qin, Xiang-Yang Li, Edward Lin, Tie-Yan Liu
Comments: Findings of EMNLP 2021
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[156] arXiv:2109.14436 (cross-list from eess.AS) [pdf, other]
Title: A Universal Deep Room Acoustics Estimator
Paula Sánchez López, Paul Callens, Milos Cernak
Comments: Room acoustics, Convolutional Recurrent Neural Network, RT60, C50, DRR, STI, SNR
Journal-ref: IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2021
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[157] arXiv:2109.14725 (cross-list from cs.LG) [pdf, other]
Title: Tiny-CRNN: Streaming Wakeword Detection In A Low Footprint Setting
Mohammad Omar Khursheed, Christin Jose, Rajath Kumar, Gengshen Fu, Brian Kulis, Santosh Kumar Cheekatmalla
Comments: arXiv admin note: substantial text overlap with arXiv:2011.12941
Subjects: Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[158] arXiv:2109.14831 (cross-list from eess.AS) [pdf, other]
Title: USEV: Universal Speaker Extraction with Visual Cue
Zexu Pan, Meng Ge, Haizhou Li
Comments: Accepted by TASLP
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[159] arXiv:2109.14992 (cross-list from cs.HC) [pdf, other]
Title: Xenakis: Experimenting with Data, Cities, and Sounds
Victor Schetinger, Ignacio Pérez-Messina, Renan Guarese, Velitchko Filipov
Comments: This manuscript heavily links to a miro board as part of an experiment in exposition, and was presented at this http URL, a workshop co-located with IEEE VIS 2021 (held virtually)
Subjects: Human-Computer Interaction (cs.HC); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[160] arXiv:2109.14994 (cross-list from eess.AS) [pdf, other]
Title: An investigation of pre-upsampling generative modelling and Generative Adversarial Networks in audio super resolution
James King, Ramon Viñas Torné, Alexander Campbell, Pietro Liò
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[161] arXiv:2109.15108 (cross-list from eess.AS) [pdf, other]
Title: Federated Learning in ASR: Not as Easy as You Think
Wentao Yu, Jan Freiwald, Sören Tewes, Fabien Huennemeyer, Dorothea Kolossa
Journal-ref: ITG Conference on Speech Communication, 2021
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[162] arXiv:2109.15127 (cross-list from eess.AS) [pdf, other]
Title: Real-Time Multi-Level Neonatal Heart and Lung Sound Quality Assessment for Telehealth Applications
Ethan Grooby, Chiranjibi Sitaula, Davood Fattahi, Reza Sameni, Kenneth Tan, Lindsay Zhou, Arrabella King, Ashwin Ramanathan, Atul Malhotra, Guy A. Dumont, Faezeh Marzbanrad
Comments: 13 pages, 8 figures, 3 tables. Paper submitted and under review in IEEE Access
Journal-ref: IEEE Access, 2022
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Signal Processing (eess.SP)
[163] arXiv:2109.15166 (cross-list from eess.AS) [pdf, other]
Title: PortaSpeech: Portable and High-Quality Generative Text-to-Speech
Yi Ren, Jinglin Liu, Zhou Zhao
Comments: Accepted by NeurIPS 2021. Source code: this https URL
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
Total of 163 entries
Showing up to 2000 entries per page: fewer | more | all
  • About
  • Help
  • contact arXivClick here to contact arXiv Contact
  • subscribe to arXiv mailingsClick here to subscribe Subscribe
  • Copyright
  • Privacy Policy
  • Web Accessibility Assistance
  • arXiv Operational Status
    Get status notifications via email or slack