Sound

Authors and titles for December 2024

Total of 231 entries : 1-25 26-50 51-75 76-100 101-125 126-150 151-175 176-200 ... 226-231

Showing up to 25 entries per page: fewer | more | all

[101] arXiv:2412.17924 [pdf, html, other]: Title: Are audio DeepFake detection models polyglots?

Bartłomiej Marek, Piotr Kawa, Piotr Syga

Comments: Keywords: Audio DeepFakes, DeepFake detection, multilingual audio DeepFakes

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[102] arXiv:2412.18061 [pdf, html, other]: Title: Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction

Hyunbae Jeon, Frederic Guintu, Rayvant Sahni

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Audio and Speech Processing (eess.AS)
[103] arXiv:2412.18157 [pdf, html, other]: Title: Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance

Yaoyun Zhang, Xuenan Xu, Mengyue Wu

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[104] arXiv:2412.18191 [pdf, html, other]: Title: Explaining Speaker and Spoof Embeddings via Probing

Xuechen Liu, Junichi Yamagishi, Md Sahidullah, Tomi kinnunen

Comments: To appear in IEEE ICASSP 2025

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[105] arXiv:2412.18217 [pdf, html, other]: Title: U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation

Shaoxiang Dang, Tetsuya Matsumoto, Yoshinori Takeuchi, Hiroaki Kudo

Journal-ref: 2024 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[106] arXiv:2412.18710 [pdf, other]: Title: Simi-SFX: A similarity-based conditioning method for controllable sound effect synthesis

Yunyi Liu, Craig Jin

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[107] arXiv:2412.18836 [pdf, html, other]: Title: MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI

Neil Shah, Ayan Kashyap, Shirish Karande, Vineet Gandhi

Comments: Accepted at IEEE ICASSP 2025

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[108] arXiv:2412.18839 [pdf, html, other]: Title: Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset

Neil Shah, Shirish Karande, Vineet Gandhi

Comments: Accepted at IEEE ICASSP 2025

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[109] arXiv:2412.18851 [pdf, html, other]: Title: Attention-Enhanced Short-Time Wiener Solution for Acoustic Echo Cancellation

Fei Zhao, Xueliang Zhang

Subjects: Sound (cs.SD)
[110] arXiv:2412.18913 [pdf, html, other]: Title: Robust Target Speaker Direction of Arrival Estimation

Zixuan Li, Shulin He, Xueliang Zhang

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[111] arXiv:2412.18955 [pdf, html, other]: Title: Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations

Julien Guinot, Elio Quinton, György Fazekas

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[112] arXiv:2412.19099 [pdf, html, other]: Title: BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement

Cunhang Fan, Enrui Liu, Andong Li, Jianhua Tao, Jian Zhou, Jiahao Li, Chengshi Zheng, Zhao Lv

Comments: Accepted by AAAI 2025

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[113] arXiv:2412.19123 [pdf, html, other]: Title: CoheDancers: Enhancing Interactive Group Dance Generation through Music-Driven Coherence Decomposition

Kaixing Yang, Xulong Tang, Haoyu Wu, Qinliang Xue, Biao Qin, Hongyan Liu, Zhaoxin Fan

Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[114] arXiv:2412.19200 [pdf, html, other]: Title: Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning

Dengming Zhang, Weitao You, Ziheng Liu, Lingyun Sun, Pei Chen

Comments: Accepted by the 39th AAAI Conference on Artificial Intelligence (AAAI-25)

Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[115] arXiv:2412.19279 [pdf, html, other]: Title: Improving Generalization for AI-Synthesized Voice Detection

Hainan Ren, Li Lin, Chun-Hao Liu, Xin Wang, Shu Hu

Comments: AAAI25

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[116] arXiv:2412.19351 [pdf, other]: Title: ETTA: Elucidating the Design Space of Text-to-Audio Models

Sang-gil Lee, Zhifeng Kong, Arushi Goel, Sungwon Kim, Rafael Valle, Bryan Catanzaro

Comments: ICML 2025. Demo: this https URL Code: this https URL

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[117] arXiv:2412.19909 [pdf, html, other]: Title: Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition

Shreya G. Upadhyay, Ali N. Salman, Carlos Busso, Chi-Chun Lee

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[118] arXiv:2412.20155 [pdf, html, other]: Title: Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting

Wooseok Han, Minki Kang, Changhun Kim, Eunho Yang

Comments: Accepted by ICASSP 2025

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[119] arXiv:2412.20914 [pdf, html, other]: Title: Language-based Audio Retrieval with Co-Attention Networks

Haoran Sun, Zimu Wang, Qiuyi Chen, Jianjun Chen, Jia Wang, Haiyang Zhang

Comments: Accepted at UIC 2024 proceedings. Accepted version

Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[120] arXiv:2412.21037 [pdf, html, other]: Title: TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Chia-Yu Hung, Navonil Majumder, Zhifeng Kong, Ambuj Mehrish, Amir Ali Bagherzadeh, Chuan Li, Rafael Valle, Bryan Catanzaro, Soujanya Poria

Comments: this https URL

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[121] arXiv:2412.00049 (cross-list from cs.MM) [pdf, html, other]: Title: A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

Luis Vilaca, Yi Yu, Paula Vinan

Comments: arXiv admin note: text overlap with arXiv:2202.13673

Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[122] arXiv:2412.00055 (cross-list from eess.AS) [pdf, html, other]: Title: High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR

Sourav Banerjee, Ayushi Agarwal, Promila Ghosh

Comments: 15 pages

Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[123] arXiv:2412.00057 (cross-list from eess.AS) [pdf, html, other]: Title: Feasibility of Mental Health Triage Call Priority Prediction Using Machine Learning

Rajib Rana, Niall Higgins, Kazi Nazmul Haque, John Reilly, Kylie Burke, Kathryn Turner, Terry Stedman

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[124] arXiv:2412.00175 (cross-list from cs.CV) [pdf, html, other]: Title: Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning

Stefan Smeu, Dragos-Alexandru Boldisor, Dan Oneata, Elisabeta Oneata

Comments: Accepted as a highlight paper at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025

Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[125] arXiv:2412.00721 (cross-list from cs.AI) [pdf, other]: Title: A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario

Zheshu Song, Ziyang Ma, Yifan Yang, Jianheng Zhuo, Xie Chen

Comments: This work hasn't been finished yet

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)

Total of 231 entries : 1-25 26-50 51-75 76-100 101-125 126-150 151-175 176-200 ... 226-231

Showing up to 25 entries per page: fewer | more | all