Skip to main content
Cornell University
We gratefully acknowledge support from the Simons Foundation, member institutions, and all contributors. Donate
arxiv logo > eess.AS

Help | Advanced Search

arXiv logo
Cornell University Logo

quick links

  • Login
  • Help Pages
  • About

Audio and Speech Processing

Authors and titles for May 2023

Total of 427 entries : 1-50 51-100 101-150 151-200 201-250 251-300 301-350 351-400 ... 401-427
Showing up to 50 entries per page: fewer | more | all
[201] arXiv:2305.05599 (cross-list from cs.SD) [pdf, other]
Title: Inter-SubNet: Speech Enhancement with Subband Interaction
Jun Chen, Wei Rao, Zilin Wang, Jiuxin Lin, Zhiyong Wu, Yannan Wang, Shidong Shang, Helen Meng
Comments: Accepted by ICASSP 2023
Subjects: Sound (cs.SD); Human-Computer Interaction (cs.HC); Audio and Speech Processing (eess.AS)
[202] arXiv:2305.05736 (cross-list from cs.SD) [pdf, other]
Title: VSMask: Defending Against Voice Synthesis Attack via Real-Time Predictive Perturbation
Yuanda Wang, Hanqing Guo, Guangjing Wang, Bocheng Chen, Qiben Yan
Subjects: Sound (cs.SD); Cryptography and Security (cs.CR); Audio and Speech Processing (eess.AS)
[203] arXiv:2305.05780 (cross-list from cs.SD) [pdf, other]
Title: Enhancing Gappy Speech Audio Signals with Generative Adversarial Networks
Deniss Strods, Alan F. Smeaton
Comments: 7 pages, 4 figures, 4 tables. 34th Irish Signals and Systems Conferences, 13-14 June 2023
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[204] arXiv:2305.06273 (cross-list from cs.CL) [pdf, other]
Title: Learning Robust Self-attention Features for Speech Emotion Recognition with Label-adaptive Mixup
Lei Kang, Lichao Zhang, Dazhi Jiang
Comments: Accepted to ICASSP 2023
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[205] arXiv:2305.06429 (cross-list from cs.SD) [pdf, other]
Title: Mispronunciation Detection of Basic Quranic Recitation Rules using Deep Learning
Ahmad Al Harere, Khloud Al Jallad
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[206] arXiv:2305.06594 (cross-list from cs.SD) [pdf, html, other]
Title: V2Meow: Meowing to the Visual Beat via Video-to-Music Generation
Kun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin, Joonseok Lee, Chris Donahue, Fei Sha, Aren Jansen, Yu Wang, Mauro Verzetti, Timo I. Denk
Comments: accepted at AAAI 2024, music samples available at this https URL
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[207] arXiv:2305.06701 (cross-list from cs.SD) [pdf, other]
Title: Extending Audio Masked Autoencoders Toward Audio Restoration
Zhi Zhong, Hao Shi, Masato Hirano, Kazuki Shimada, Kazuya Tateishi, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji
Comments: WASPAA this http URL 2023 this http URL use of this material is this http URL from IEEE must be obtained for all other uses,in any current or future media,including reprinting/republishing this material for advertising or promotional purposes, creating new collective works,for resale or redistribution to servers or lists,or reuse of any copyrighted component of this work in other works
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[208] arXiv:2305.06806 (cross-list from cs.SD) [pdf, other]
Title: HappyQuokka System for ICASSP 2023 Auditory EEG Challenge
Zhenyu Piao, Miseul Kim, Hyungchan Yoon, Hong-Goo Kang
Comments: First Place in Task 2 of Auditory EEG decoding Challenge, which is part of ICASSP Signal Processing Grand Challenge (SPGC) 2023
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[209] arXiv:2305.06908 (cross-list from cs.SD) [pdf, other]
Title: CoMoSpeech: One-Step Speech and Singing Voice Synthesis via Consistency Model
Zhen Ye, Wei Xue, Xu Tan, Jie Chen, Qifeng Liu, Yike Guo
Comments: Accepted to ACM MM 2023
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[210] arXiv:2305.07132 (cross-list from cs.SD) [pdf, other]
Title: Tackling Interpretability in Audio Classification Networks with Non-negative Matrix Factorization
Jayneel Parekh, Sanjeel Parekh, Pavlo Mozharovskyi, Gaël Richard, Florence d'Alché-Buc
Comments: Under submission at IEEE/ACM TASLP. arXiv admin note: text overlap with arXiv:2202.11479
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[211] arXiv:2305.07216 (cross-list from cs.LG) [pdf, html, other]
Title: Versatile audio-visual learning for emotion recognition
Lucas Goncalves, Seong-Gyun Leem, Wei-Cheng Lin, Berrak Sisman, Carlos Busso
Comments: 18 pages, 4 Figures, 3 tables (published at IEEE Transactions on Affective Computing)
Subjects: Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[212] arXiv:2305.07223 (cross-list from cs.SD) [pdf, html, other]
Title: Transavs: End-To-End Audio-Visual Segmentation With Transformer
Yuhang Ling, Yuxi Li, Zhenye Gan, Jiangning Zhang, Mingmin Chi, Yabiao Wang
Comments: 4 pages, 3 figures
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[213] arXiv:2305.07243 (cross-list from cs.SD) [pdf, other]
Title: Better speech synthesis through scaling
James Betker
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[214] arXiv:2305.07347 (cross-list from cs.SD) [pdf, other]
Title: Music Rearrangement Using Hierarchical Segmentation
Christos Plachouras, Marius Miron
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[215] arXiv:2305.07389 (cross-list from cs.CL) [pdf, other]
Title: Investigating the Sensitivity of Automatic Speech Recognition Systems to Phonetic Variation in L2 Englishes
Emma O'Neill, Julie Carson-Berndsen
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[216] arXiv:2305.07447 (cross-list from cs.SD) [pdf, other]
Title: Universal Source Separation with Weakly Labelled Data
Qiuqiang Kong, Ke Chen, Haohe Liu, Xingjian Du, Taylor Berg-Kirkpatrick, Shlomo Dubnov, Mark D. Plumbley
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[217] arXiv:2305.07455 (cross-list from cs.CL) [pdf, other]
Title: Improving Cascaded Unsupervised Speech Translation with Denoising Back-translation
Yu-Kuan Fu, Liang-Hsuan Tseng, Jiatong Shi, Chen-An Li, Tsu-Yuan Hsu, Shinji Watanabe, Hung-yi Lee
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[218] arXiv:2305.07489 (cross-list from cs.SD) [pdf, html, other]
Title: Benchmarks and leaderboards for sound demixing tasks
Roman Solovyev, Alexander Stempkovskiy, Tatiana Habruseva
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[219] arXiv:2305.07499 (cross-list from cs.SD) [pdf, other]
Title: Device-Robust Acoustic Scene Classification via Impulse Response Augmentation
Tobias Morocutti, Florian Schmid, Khaled Koutini, Gerhard Widmer
Comments: In Proceedings of the 31st European Signal Processing Conference, EUSIPCO 2023. Source Code available at: this https URL
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[220] arXiv:2305.07828 (cross-list from cs.SD) [pdf, other]
Title: Description and Discussion on DCASE 2023 Challenge Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
Kota Dohi, Keisuke Imoto, Noboru Harada, Daisuke Niizumi, Yuma Koizumi, Tomoya Nishida, Harsh Purohit, Ryo Tanabe, Takashi Endo, Yohei Kawaguchi
Comments: anomaly detection, acoustic condition monitoring, domain shift, first-shot problem, DCASE Challenge, Accepted in DCASE2023 Workshop
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[221] arXiv:2305.07909 (cross-list from cs.SD) [pdf, other]
Title: Higher-Order Frequency Modulation Synthesis
Victor Lazzarini, Joseph Timoney
Comments: 15 pages, 6 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[222] arXiv:2305.07952 (cross-list from cs.SD) [pdf, other]
Title: APNet: An All-Frame-Level Neural Vocoder Incorporating Direct Prediction of Amplitude and Phase Spectra
Yang Ai, Zhen-Hua Ling
Comments: Accepted by IEEE/ACM Transactions on Audio, Speech, and Language Processing. Codes are available
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[223] arXiv:2305.07960 (cross-list from cs.SD) [pdf, other]
Title: Sound-to-Vibration Transformation for Sensorless Motor Health Monitoring
Ozer Can Devecioglu, Serkan Kiranyaz, Amer Elhmes, Sadok Sassi, Turker Ince, Onur Avci, Mohammad Hesam Soleimani-Babakamali, Ertugrul Taciroglu, Moncef Gabbouj
Subjects: Sound (cs.SD); Human-Computer Interaction (cs.HC); Audio and Speech Processing (eess.AS)
[224] arXiv:2305.08014 (cross-list from cs.CV) [pdf, other]
Title: Surface EMG-Based Inter-Session/Inter-Subject Gesture Recognition by Leveraging Lightweight All-ConvNet and Transfer Learning
Md. Rabiul Islam, Daniel Massicotte, Philippe Y. Massicotte, Wei-Ping Zhu
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[225] arXiv:2305.08029 (cross-list from cs.SD) [pdf, html, other]
Title: REMAST: Real-time Emotion-based Music Arrangement with Soft Transition
Zihao Wang, Le Ma, Chen Zhang, Bo Han, Yunfei Xu, Yikai Wang, Xinyi Chen, HaoRong Hong, Wenbo Liu, Xinda Wu, Kejun Zhang
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[226] arXiv:2305.08067 (cross-list from cs.CL) [pdf, other]
Title: Improving End-to-End SLU performance with Prosodic Attention and Distillation
Shangeth Rajaa
Comments: Submitted to InterSpeech 2023
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[227] arXiv:2305.08099 (cross-list from cs.SD) [pdf, other]
Title: Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech Representations
Weiwei Lin, Chenhang He, Man-Wai Mak, Youzhi Tu
Comments: accepted by ICML 2023
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[228] arXiv:2305.08292 (cross-list from cs.SD) [pdf, other]
Title: ForkNet: Simultaneous Time and Time-Frequency Domain Modeling for Speech Enhancement
Feng Dang, Qi Hu, Pengyuan Zhang, Yonghong Yan
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[229] arXiv:2305.08541 (cross-list from cs.SD) [pdf, other]
Title: Ripple sparse self-attention for monaural speech enhancement
Qiquan Zhang, Hongxu Zhu, Qi Song, Xinyuan Qian, Zhaoheng Ni, Haizhou Li
Comments: 5 pages, ICASSP 2023 published
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[230] arXiv:2305.08706 (cross-list from cs.CL) [pdf, other]
Title: Understanding and Bridging the Modality Gap for Speech Translation
Qingkai Fang, Yang Feng
Comments: ACL 2023 main conference
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[231] arXiv:2305.08709 (cross-list from cs.CL) [pdf, other]
Title: Back Translation for Speech-to-text Translation Without Transcripts
Qingkai Fang, Yang Feng
Comments: ACL 2023 main conference
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[232] arXiv:2305.09167 (cross-list from cs.SD) [pdf, other]
Title: Adversarial Speaker Disentanglement Using Unannotated External Data for Self-supervised Representation Based Voice Conversion
Xintao Zhao, Shuai Wang, Yang Chao, Zhiyong Wu, Helen Meng,
Comments: Accepted by ICME 2023
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[233] arXiv:2305.09302 (cross-list from cs.CV) [pdf, other]
Title: Pink-Eggs Dataset V1: A Step Toward Invasive Species Management Using Deep Learning Embedded Solutions
Di Xu, Yang Zhao, Xiang Hao, Xin Meng
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[234] arXiv:2305.09463 (cross-list from cs.SD) [pdf, other]
Title: Low-complexity deep learning frameworks for acoustic scene classification using teacher-student scheme and multiple spectrograms
Lam Pham, Dat Ngo, Cam Le, Anahid Jalali, Alexander Schindler
Comments: arXiv admin note: text overlap with arXiv:2206.06057
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[235] arXiv:2305.09489 (cross-list from cs.SD) [pdf, other]
Title: Discrete Diffusion Probabilistic Models for Symbolic Music Generation
Matthias Plasser, Silvan Peter, Gerhard Widmer
Comments: In Proceedings of the 32nd International Joint Conference on Artificial Intelligence (IJCAI-23), Macau, China
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[236] arXiv:2305.09559 (cross-list from cs.SD) [pdf, other]
Title: Robust and lightweight audio fingerprint for Automatic Content Recognition
Anoubhav Agarwaal, Prabhat Kanaujia, Sartaki Sinha Roy, Susmita Ghose
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[237] arXiv:2305.09636 (cross-list from cs.SD) [pdf, other]
Title: SoundStorm: Efficient Parallel Audio Generation
Zalán Borsos, Matt Sharifi, Damien Vincent, Eugene Kharitonov, Neil Zeghidour, Marco Tagliasacchi
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[238] arXiv:2305.09652 (cross-list from cs.CL) [pdf, other]
Title: The Interpreter Understands Your Meaning: End-to-end Spoken Language Understanding Aided by Speech Translation
Mutian He, Philip N. Garner
Comments: 16 pages, 3 figures; accepted by Findings of EMNLP 2023
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[239] arXiv:2305.09690 (cross-list from cs.SD) [pdf, other]
Title: A Whisper transformer for audio captioning trained with synthetic captions and transfer learning
Marek Kadlčík, Adam Hájek, Jürgen Kieslich, Radosław Winiecki
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[240] arXiv:2305.09764 (cross-list from cs.CL) [pdf, other]
Title: Application-Agnostic Language Modeling for On-Device ASR
Markus Nußbaum-Thom, Lyan Verwimp, Youssef Oualil
Comments: accepted for ACL 2023 industry track
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[241] arXiv:2305.10270 (cross-list from cs.CL) [pdf, other]
Title: Boosting Local Spectro-Temporal Features for Speech Analysis
Michael Guerzhoy
Comments: Master's project, University of Toronto, 2010
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[242] arXiv:2305.10321 (cross-list from cs.CL) [pdf, other]
Title: Controllable Speaking Styles Using a Large Language Model
Atli Thor Sigurgeirsson, Simon King
Comments: Submitted to ICASSP 2024
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[243] arXiv:2305.10358 (cross-list from cs.CR) [pdf, other]
Title: NUANCE: Near Ultrasound Attack On Networked Communication Environments
Forrest McKee, David Noever
Subjects: Cryptography and Security (cs.CR); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[244] arXiv:2305.10615 (cross-list from cs.SD) [pdf, html, other]
Title: ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
Jiatong Shi, Dan Berrebbi, William Chen, Ho-Lam Chung, En-Pei Hu, Wei Ping Huang, Xuankai Chang, Shang-Wen Li, Abdelrahman Mohamed, Hung-yi Lee, Shinji Watanabe
Comments: Accepted by Interspeech
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[245] arXiv:2305.10649 (cross-list from cs.SD) [pdf, other]
Title: ZeroPrompt: Streaming Acoustic Encoders are Zero-Shot Masked LMs
Xingchen Song, Di Wu, Binbin Zhang, Zhendong Peng, Bo Dang, Fuping Pan, Zhiyong Wu
Comments: accepted by interspeech 2023
Journal-ref: @inproceedings{song23c_interspeech, year=2023, booktitle={Proc. INTERSPEECH 2023}, pages={1648--1652}}
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[246] arXiv:2305.10652 (cross-list from cs.SD) [pdf, html, other]
Title: Speech Separation based on Contrastive Learning and Deep Modularization
Peter Ochieng
Comments: arXiv admin note: substantial text overlap with arXiv:2212.00369
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[247] arXiv:2305.10666 (cross-list from cs.CL) [pdf, html, other]
Title: A unified front-end framework for English text-to-speech synthesis
Zelin Ying, Chen Li, Yu Dong, Qiuqiang Kong, Qiao Tian, Yuanyuan Huo, Yuxuan Wang
Comments: Accepted in ICASSP 2024
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[248] arXiv:2305.10680 (cross-list from cs.SD) [pdf, other]
Title: Accurate and Reliable Confidence Estimation Based on Non-Autoregressive End-to-End Speech Recognition System
Xian Shi, Haoneng Luo, Zhifu Gao, Shiliang Zhang, Zhijie Yan
Comments: 5 pages, 4 figures, Interspeech2023
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[249] arXiv:2305.10686 (cross-list from cs.SD) [pdf, other]
Title: RMSSinger: Realistic-Music-Score based Singing Voice Synthesis
Jinzheng He, Jinglin Liu, Zhenhui Ye, Rongjie Huang, Chenye Cui, Huadai Liu, Zhou Zhao
Comments: Accepted by Finding of ACL2023
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[250] arXiv:2305.10704 (cross-list from cs.SD) [pdf, other]
Title: Attention-based Encoder-Decoder Network for End-to-End Neural Speaker Diarization with Target Speaker Attractor
Zhengyang Chen, Bing Han, Shuai Wang, Yanmin Qian
Comments: Accepted by InterSpeech 2023
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
Total of 427 entries : 1-50 51-100 101-150 151-200 201-250 251-300 301-350 351-400 ... 401-427
Showing up to 50 entries per page: fewer | more | all
  • About
  • Help
  • contact arXivClick here to contact arXiv Contact
  • subscribe to arXiv mailingsClick here to subscribe Subscribe
  • Copyright
  • Privacy Policy
  • Web Accessibility Assistance
  • arXiv Operational Status
    Get status notifications via email or slack