Audio and Speech Processing

Authors and titles for March 2024

Total of 213 entries : 1-25 26-50 51-75 76-100 101-125 126-150 151-175 176-200 ... 201-213

Showing up to 25 entries per page: fewer | more | all

[101] arXiv:2403.03762 (cross-list from eess.SP) [pdf, html, other]: Title: Room Impulse Response Estimation using Optimal Transport: Simulation-Informed Inference

David Sundström, Anton Björkman, Andreas Jakobsson, Filip Elvander

Subjects: Signal Processing (eess.SP); Audio and Speech Processing (eess.AS)
[102] arXiv:2403.03947 (cross-list from cs.SD) [pdf, html, other]: Title: Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset

Pedro Ramoneda, Minhee Lee, Dasaem Jeong, J.J. Valero-Mas, Xavier Serra

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[103] arXiv:2403.04111 (cross-list from cs.SD) [pdf, other]: Title: Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication

Yejin Jeon, Gary Geunbae Lee

Comments: Accepted to EACL Main 2024

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[104] arXiv:2403.04178 (cross-list from cs.CL) [pdf, html, other]: Title: Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation

Sai Akarsh, Vamshi Raghusimha, Anindita Mondal, Anil Vuppala

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[105] arXiv:2403.04245 (cross-list from cs.SD) [pdf, html, other]: Title: A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition

Yusheng Dai, Hang Chen, Jun Du, Ruoyu Wang, Shihao Chen, Jiefeng Ma, Haotian Wang, Chin-Hui Lee

Comments: the paper is accepted by CVPR2024

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[106] arXiv:2403.04594 (cross-list from cs.SD) [pdf, html, other]: Title: A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds

Xuenan Xu, Xiaohang Xu, Zeyu Xie, Pingyue Zhang, Mengyue Wu, Kai Yu

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[107] arXiv:2403.04654 (cross-list from cs.CV) [pdf, html, other]: Title: Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention

R. Gnana Praveen, Jahangir Alam

Comments: Accepted to FG2024

Subjects: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[108] arXiv:2403.04661 (cross-list from cs.CV) [pdf, html, other]: Title: Dynamic Cross Attention for Audio-Visual Person Verification

R. Gnana Praveen, Jahangir Alam

Comments: Accepted to FG2024

Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[109] arXiv:2403.05010 (cross-list from cs.SD) [pdf, html, other]: Title: RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction

Peng Liu, Dongyang Dai, Zhiyong Wu

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[110] arXiv:2403.05380 (cross-list from cs.SD) [pdf, html, other]: Title: Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings

Mahyar Gohari, Paolo Bestagini, Sergio Benini, Nicola Adami

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[111] arXiv:2403.05583 (cross-list from cs.HC) [pdf, html, other]: Title: A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition

Tyler Benster, Guy Wilson, Reshef Elisha, Francis R Willett, Shaul Druckmann

Subjects: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[112] arXiv:2403.05772 (cross-list from cs.SD) [pdf, html, other]: Title: sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks

Qu Yang, Qianhui Liu, Nan Li, Meng Ge, Zeyang Song, Haizhou Li

Comments: Accepted by ICASSP 2024

Subjects: Sound (cs.SD); Neural and Evolutionary Computing (cs.NE); Audio and Speech Processing (eess.AS)
[113] arXiv:2403.05820 (cross-list from cs.SD) [pdf, html, other]: Title: An Audio-textual Diffusion Model For Converting Speech Signals Into Ultrasound Tongue Imaging Data

Yudong Yang, Rongfeng Su, Xiaokang Liu, Nan Yan, Lan Wang

Comments: ICASSP2024 Accept

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[114] arXiv:2403.05834 (cross-list from cs.MM) [pdf, html, other]: Title: Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information

Qiaochu Huang, Xu He, Boshi Tang, Haolin Zhuang, Liyang Chen, Shuochen Gao, Zhiyong Wu, Haozhi Huang, Helen Meng

Subjects: Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[115] arXiv:2403.05989 (cross-list from cs.SD) [pdf, html, other]: Title: HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling

Chunhui Wang, Chang Zeng, Bowen Zhang, Ziyang Ma, Yefan Zhu, Zifeng Cai, Jian Zhao, Zhonglin Jiang, Yong Chen

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[116] arXiv:2403.06100 (cross-list from cs.HC) [pdf, html, other]: Title: Automatic design optimization of preference-based subjective evaluation with online learning in crowdsourcing environment

Yusuke Yasuda, Tomoki Toda

Subjects: Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
[117] arXiv:2403.06260 (cross-list from cs.CL) [pdf, html, other]: Title: SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations

Amit Meghanani, Thomas Hain

Comments: Accepted at ICASSP 2024

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[118] arXiv:2403.06387 (cross-list from cs.SD) [pdf, html, other]: Title: Towards Decoupling Frontend Enhancement and Backend Recognition in Monaural Robust ASR

Yufeng Yang, Ashutosh Pandey, DeLiang Wang

Comments: Submitted to IEEE/ACM Transactions on Audio, Speech and Language Processing. arXiv admin note: text overlap with arXiv:2210.13318

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[119] arXiv:2403.06404 (cross-list from cs.SD) [pdf, html, other]: Title: Cosine Scoring with Uncertainty for Neural Speaker Embedding

Qiongqiong Wang, Kong Aik Lee

Comments: 5 pages, 4 figures

Journal-ref: IEEE Signal Processing Letters 2024

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[120] arXiv:2403.06487 (cross-list from cs.CL) [pdf, html, other]: Title: Multilingual Turn-taking Prediction Using Voice Activity Projection

Koji Inoue, Bing'er Jiang, Erik Ekstedt, Tatsuya Kawahara, Gabriel Skantze

Comments: This paper has been accepted for presentation at The 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) and represents the author's version of the work

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[121] arXiv:2403.07675 (cross-list from cs.SD) [pdf, html, other]: Title: Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers

Changsheng Quan, Xiaofei Li

Comments: Accepted by IEEE Signal Processing Letters

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[122] arXiv:2403.07802 (cross-list from cs.SD) [pdf, html, other]: Title: Boosting keyword spotting through on-device learnable user speech characteristics

Cristian Cioflan, Lukas Cavigelli, Luca Benini

Comments: 5 pages, 3 tables, 2 figures. Accepted as a full paper by the tinyML Research Symposium 2024

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[123] arXiv:2403.07938 (cross-list from cs.SD) [pdf, html, other]: Title: Text-to-Audio Generation Synchronized with Videos

Shentong Mo, Jing Shi, Yapeng Tian

Comments: arXiv admin note: text overlap with arXiv:2305.12903

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[124] arXiv:2403.07995 (cross-list from cs.SD) [pdf, html, other]: Title: Motifs, Phrases, and Beyond: The Modelling of Structure in Symbolic Music Generation

Keshav Bhandari, Simon Colton

Comments: Accepted to 13th International Conference on Artificial Intelligence in Music, Sound, Art and Design (EvoMUSART) 2024

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Symbolic Computation (cs.SC); Audio and Speech Processing (eess.AS)
[125] arXiv:2403.08164 (cross-list from cs.SD) [pdf, html, other]: Title: EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech

Ziqi Liang, Haoxiang Shi, Jiawei Wang, Keda Lu

Comments: Accepted by the 27th IEEE International Conference on Computer Supported Cooperative Work in Design (IEEE CSCWD 2024). arXiv admin note: substantial text overlap with arXiv:2211.01948

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)

Total of 213 entries : 1-25 26-50 51-75 76-100 101-125 126-150 151-175 176-200 ... 201-213

Showing up to 25 entries per page: fewer | more | all