Skip to main content
Cornell University
We gratefully acknowledge support from the Simons Foundation, member institutions, and all contributors. Donate
arxiv logo > cs.SD

Help | Advanced Search

arXiv logo
Cornell University Logo

quick links

  • Login
  • Help Pages
  • About

Sound

Authors and titles for December 2024

Total of 231 entries : 26-125 101-200 201-231
Showing up to 100 entries per page: fewer | more | all
[26] arXiv:2412.05123 [pdf, html, other]
Title: Applying Automatic Differentiation to Optimize Differential Microphone Array Designs
Siminfar Samakoush Galougah, Ramani Duraiswami
Comments: 6 pages, 9 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[27] arXiv:2412.05436 [pdf, html, other]
Title: pyAMPACT: A Score-Audio Alignment Toolkit for Performance Data Estimation and Multi-modal Processing
Johanna Devaney, Daniel McKemie, Alex Morgan
Comments: International Society for Music Information Retrieval, Late Breaking Demo
Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[28] arXiv:2412.05558 [pdf, html, other]
Title: WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
Feng Li, Jiusong Luo, Wanjun Xia
Comments: Accepted by 31st International Conference on MultiMedia Modeling (MMM2025)
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[29] arXiv:2412.05951 [pdf, html, other]
Title: When Vision Models Meet Parameter Efficient Look-Aside Adapters Without Large-Scale Audio Pretraining
Juan Yeo, Jinkwan Jang, Kyubyung Chae, Seongkyu Mun, Taesup Kim
Comments: 5 pages, 3 figures
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[30] arXiv:2412.06001 [pdf, html, other]
Title: M6: Multi-generator, Multi-domain, Multi-lingual and cultural, Multi-genres, Multi-instrument Machine-Generated Music Detection Databases
Yupei Li, Hanqian Li, Lucia Specia, Björn W. Schuller
Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[31] arXiv:2412.06208 [pdf, html, other]
Title: Pilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization
Fei Yu, Zhe Xiang, Nan Che, Zhuoran Zhang, Yuandi Li, Junxiao Xue, Zhiguo Wan
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[32] arXiv:2412.06296 [pdf, html, other]
Title: VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features
Sifei Li, Binxin Yang, Chunji Yin, Chong Sun, Yuxin Zhang, Weiming Dong, Chen Li
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[33] arXiv:2412.06581 [pdf, other]
Title: EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
Weizhen Bian, Yubo Zhou, Kaitai Zhang, Xiaohan Gu
Comments: I did not obtain the necessary approval from my academic supervisor prior to submission and there are issues with my current paper
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[34] arXiv:2412.06617 [pdf, html, other]
Title: AI TrackMate: Finally, Someone Who Will Give Your Music More Than Just "Sounds Great!"
Yi-Lin Jiang, Chia-Ho Hsiung, Yen-Tung Yeh, Lu-Rong Chen, Bo-Yu Chen
Comments: Accepted for the NeurIPS 2024 Creative AI Track
Subjects: Sound (cs.SD); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[35] arXiv:2412.06660 [pdf, html, other]
Title: MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
Shansong Liu, Atin Sakkeer Hussain, Qilong Wu, Chenshuo Sun, Ying Shan
Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[36] arXiv:2412.06703 [pdf, html, other]
Title: Source Separation & Automatic Transcription for Music
Bradford Derby, Lucas Dunker, Samarth Galchar, Shashank Jarmale, Akash Setti
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[37] arXiv:2412.06965 [pdf, html, other]
Title: Improving Source Extraction with Diffusion and Consistency Models
Tornike Karchkhadze, Mohammad Rasool Izadi, Shuo Zhang
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[38] arXiv:2412.07316 [pdf, html, other]
Title: Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
Rui Zhou, Akinori Ito, Takashi Nose
Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[39] arXiv:2412.07948 [pdf, html, other]
Title: Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
Jan Retkowski, Jakub Stępniak, Mateusz Modrzejewski
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[40] arXiv:2412.08112 [pdf, html, other]
Title: Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration
Haowei Lou, Helen Paik, Wen Hu, Lina Yao
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[41] arXiv:2412.08117 [pdf, html, other]
Title: LatentSpeech: Latent Diffusion for Text-To-Speech Generation
Haowei Lou, Helen Paik, Pari Delir Haghighi, Wen Hu, Lina Yao
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[42] arXiv:2412.08237 [pdf, html, other]
Title: TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
Xingchen Song, Mengtao Xing, Changwei Ma, Shengqiang Li, Di Wu, Binbin Zhang, Fuping Pan, Dinghao Zhou, Yuekai Zhang, Shun Lei, Zhendong Peng, Zhiyong Wu
Comments: Technical Report
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[43] arXiv:2412.08247 [pdf, html, other]
Title: MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
Junjie Li, Ke Zhang, Shuai Wang, Kong Aik Lee, Man-Wai Mak, Haizhou Li
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[44] arXiv:2412.08312 [pdf, html, other]
Title: A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
Sowmya Cheripally
Comments: 7 pages, 5 figures, 2 tables
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[45] arXiv:2412.08356 [pdf, html, other]
Title: Zero-Shot Mono-to-Binaural Speech Synthesis
Alon Levkovitch, Julian Salazar, Soroosh Mariooryad, RJ Skerry-Ryan, Nadav Bar, Bastiaan Kleijn, Eliya Nachmani
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[46] arXiv:2412.08504 [pdf, html, other]
Title: PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
Yifan Xie, Tao Feng, Xin Zhang, Xiangyang Luo, Zixuan Guo, Weijiang Yu, Heng Chang, Fei Ma, Fei Richard Yu
Comments: 9 pages, accepted by AAAI 2025
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Graphics (cs.GR); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[47] arXiv:2412.08550 [pdf, html, other]
Title: Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
Hugo Flores García, Oriol Nieto, Justin Salamon, Bryan Pardo, Prem Seetharaman
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[48] arXiv:2412.08577 [pdf, html, other]
Title: Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
Hongming Guo, Ruibo Fu, Yizhong Geng, Shuai Liu, Shuchen Shi, Tao Wang, Chunyu Qiang, Chenxing Li, Ya Li, Zhengqi Wen, Yukun Liu, Xuefei Liu
Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[49] arXiv:2412.08608 [pdf, html, other]
Title: AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models
Mintong Kang, Chejian Xu, Bo Li
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Audio and Speech Processing (eess.AS)
[50] arXiv:2412.08683 [pdf, html, other]
Title: Emotional Vietnamese Speech-Based Depression Diagnosis Using Dynamic Attention Mechanism
Quang-Anh N.D., Manh-Hung Ha, Thai Kim Dinh, Minh-Duc Pham, Ninh Nguyen Van
Comments: 9 Page, 5 Figures
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[51] arXiv:2412.08856 [pdf, html, other]
Title: Complex-Cycle-Consistent Diffusion Model for Monaural Speech Enhancement
Yi Li, Yang Sun, Plamen Angelov
Comments: AAAI 2025
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[52] arXiv:2412.08944 [pdf, html, other]
Title: Interpreting Graphic Notation with MusicLDM: An AI Improvisation of Cornelius Cardew's Treatise
Tornike Karchkhadze, Keren Shao, Shlomo Dubnov
Journal-ref: 2024 IEEE International Conference on Big Data (Big Data)
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[53] arXiv:2412.08988 [pdf, html, other]
Title: EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
Gaoxiang Cong, Jiadong Pan, Liang Li, Yuankai Qi, Yuxin Peng, Anton van den Hengel, Jian Yang, Qingming Huang
Comments: Accepted to CVPR 2025
Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[54] arXiv:2412.09032 [pdf, html, other]
Title: Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
Zhoulin Ji, Chenhao Lin, Hang Wang, Chao Shen
Comments: IJCAI 2024
Journal-ref: Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24, 2024, pp. 413-421
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[55] arXiv:2412.09168 [pdf, html, other]
Title: YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
Zihao Chen, Haomin Zhang, Xinhan Di, Haoyu Wang, Sizhe Shan, Junjie Zheng, Yunming Liang, Yihan Fan, Xinfa Zhu, Wenjie Tian, Yihua Wang, Chaofan Ding, Lei Xie
Comments: 16 pages, 4 figures
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[56] arXiv:2412.09195 [pdf, html, other]
Title: On the Generation and Removal of Speaker Adversarial Perturbation for Voice-Privacy Protection
Chenyang Guo, Liping Chen, Zhuhai Li, Kong Aik Lee, Zhen-Hua Ling, Wu Guo
Comments: 6 pages, 3 figures, published to IEEE SLT Workshop 2024
Journal-ref: 2024 IEEE Spoken Language Technology Workshop (SLT), 2024, pp. 1197-1202
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[57] arXiv:2412.09317 [pdf, html, other]
Title: Multimodal Sentiment Analysis based on Video and Audio Inputs
Antonio Fernandez, Suzan Awinat
Comments: Presented as a full paper in the 15th International Conference on Emerging Ubiquitous Systems and Pervasive Networks (EUSPN 2024) October 28-30, 2024, Leuven, Belgium
Journal-ref: Procedia Computer Science, Volume 251, 2024, Pages 41-48, ISSN 1877-0509
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[58] arXiv:2412.09467 [pdf, html, other]
Title: Audios Don't Lie: Multi-Frequency Channel Attention Mechanism for Audio Deepfake Detection
Yangguang Feng
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[59] arXiv:2412.09789 [pdf, html, other]
Title: SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
Sonal Kumar, Prem Seetharaman, Justin Salamon, Dinesh Manocha, Oriol Nieto
Comments: Website: this https URL
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[60] arXiv:2412.09928 [pdf, html, other]
Title: Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's Disease Identification
Yifan Gao, Long Guo, Hong Liu
Comments: ICASSP 2025
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[61] arXiv:2412.10011 [pdf, html, other]
Title: Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
Niloy Kumar Kundu, Sarah Kobir, Md. Rayhan Ahmed, Tahmina Aktar, Niloya Roy
Comments: 42 pages,10 figures
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[62] arXiv:2412.10117 [pdf, html, other]
Title: CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Zhihao Du, Yuxuan Wang, Qian Chen, Xian Shi, Xiang Lv, Tianyu Zhao, Zhifu Gao, Yexin Yang, Changfeng Gao, Hui Wang, Fan Yu, Huadai Liu, Zhengyan Sheng, Yue Gu, Chong Deng, Wen Wang, Shiliang Zhang, Zhijie Yan, Jingren Zhou
Comments: Tech report, work in progress
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[63] arXiv:2412.10469 [pdf, other]
Title: Comparative Analysis of Mel-Frequency Cepstral Coefficients and Wavelet Based Audio Signal Processing for Emotion Detection and Mental Health Assessment in Spoken Speech
Idoko Agbo, Dr Hoda El-Sayed, M.D Kamruzzan Sarker
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[64] arXiv:2412.10481 [pdf, other]
Title: Tipping Points, Pulse Elasticity and Tonal Tension: An Empirical Study on What Generates Tipping Points
Canishk Naik (CAM, LSE), Elaine Chew (Repmus, CNRS, STMS)
Comments: International Society for Music Information Retrieval Conference, Oct 2017, Suzhou, China, China
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Audio and Speech Processing (eess.AS)
[65] arXiv:2412.10649 [pdf, html, other]
Title: Hidden Echoes Survive Training in Audio To Audio Generative Instrument Models
Christopher J. Tralie, Matt Amery, Benjamin Douglas, Ian Utz
Comments: 8 pages, 11 Figures, Proceedings of 2025 AAAI Workshop on AI for Music
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[66] arXiv:2412.10792 [pdf, html, other]
Title: Audio-based Anomaly Detection in Industrial Machines Using Deep One-Class Support Vector Data Description
Sertac Kilickaya, Mete Ahishali, Cansu Celebioglu, Fahad Sohrab, Levent Eren, Turker Ince, Murat Askar, Moncef Gabbouj
Comments: To be published in 2025 IEEE Symposium Series on Computational Intelligence
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[67] arXiv:2412.10857 [pdf, html, other]
Title: Robust Persian Digit Recognition in Noisy Environments Using Hybrid CNN-BiGRU Model
Ali Nasr-Esfahani, Mehdi Bekrani, Roozbeh Rajabi
Comments: 6 pages, two columns, submitted to Pattern Recognition Letters
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[68] arXiv:2412.10968 [pdf, html, other]
Title: Composers' Evaluations of an AI Music Tool: Insights for Human-Centred Design
Eleanor Row, György Fazekas
Comments: Accepted to NeurIPS 2024 Workshop on Generative AI and Creativity: A dialogue between machine learning researchers and creative professionals in Vancouver, Canada
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[69] arXiv:2412.11272 [pdf, html, other]
Title: WhisperFlow: speech foundation models in real time
Rongxiang Wang, Zhiming Xu, Felix Xiaozhu Lin
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[70] arXiv:2412.11449 [pdf, html, other]
Title: Whisper-GPT: A Hybrid Representation Audio Large Language Model
Prateek Verma
Comments: 6 pages, 3 figures. 50th International Conference on Acoustics, Speech and Signal Processing, Hyderabad, India
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[71] arXiv:2412.11551 [pdf, html, other]
Title: Region-Based Optimization in Continual Learning for Audio Deepfake Detection
Yujie Chen, Jiangyan Yi, Cunhang Fan, Jianhua Tao, Yong Ren, Siding Zeng, Chu Yuan Zhang, Xinrui Yan, Hao Gu, Jun Xue, Chenglong Wang, Zhao Lv, Xiaohui Zhang
Comments: Accepted by AAAI 2025
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[72] arXiv:2412.11769 [pdf, other]
Title: Does it Chug? Towards a Data-Driven Understanding of Guitar Tone Description
Pratik Sutar, Jason Naradowsky, Yusuke Miyao
Comments: Accepted for publication at the 3rd Workshop on NLP for Music and Audio (NLP4MusA 2024)
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[73] arXiv:2412.11907 [pdf, html, other]
Title: AudioCIL: A Python Toolbox for Audio Class-Incremental Learning with Multiple Scenes
Qisheng Xu, Yulin Sun, Yi Su, Qian Zhu, Xiaoyi Tan, Hongyu Wen, Zijian Gao, Kele Xu, Yong Dou, Dawei Feng
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[74] arXiv:2412.11943 [pdf, html, other]
Title: autrainer: A Modular and Extensible Deep Learning Toolkit for Computer Audition Tasks
Simon Rampp, Andreas Triantafyllopoulos, Manuel Milling, Björn W. Schuller
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[75] arXiv:2412.12111 [pdf, html, other]
Title: Voice Biomarker Analysis and Automated Severity Classification of Dysarthric Speech in a Multilingual Context
Eunjung Yeo
Comments: SNU Doctoral thesis
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[76] arXiv:2412.12395 [pdf, html, other]
Title: Sound Classification of Four Insect Classes
Yinxuan Wang, Sudip Vhaduri
Comments: The manuscript is in submission
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[77] arXiv:2412.12498 [pdf, html, other]
Title: Hierarchical Control of Emotion Rendering in Speech Synthesis
Sho Inoue, Kun Zhou, Shuai Wang, Haizhou Li
Comments: Accepted to IEEE Transactions on Affective Computing
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[78] arXiv:2412.12512 [pdf, html, other]
Title: Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
Yun Liu, Xuechen Liu, Xiaoxiao Miao, Junichi Yamagishi
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[79] arXiv:2412.12619 [pdf, html, other]
Title: Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes
Kuiyuan Zhang, Zhongyun Hua, Rushi Lan, Yushu Zhang, Yifang Guo
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[80] arXiv:2412.12760 [pdf, html, other]
Title: CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
He Wang, Xucheng Wan, Naijun Zheng, Kai Liu, Huan Zhou, Guojian Li, Lei Xie
Comments: Accepted by ICASSP 2025. 5 pages, 2 figures
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[81] arXiv:2412.13037 [pdf, html, other]
Title: TAME: Temporal Audio-based Mamba for Enhanced Drone Trajectory Estimation and Classification
Zhenyuan Xiao, Huanran Hu, Guili Xu, Junwei He
Comments: This paper has been accepted for presentation at the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2025. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[82] arXiv:2412.13279 [pdf, html, other]
Title: Synthetic Speech Classification: IEEE Signal Processing Cup 2022 challenge
Mahieyin Rahmun, Rafat Hasan Khan, Tanjim Taharat Aurpa, Sadia Khan, Zulker Nayeen Nahiyan, Mir Sayad Bin Almas, Rakibul Hasan Rajib, Syeda Sakira Hassan
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[83] arXiv:2412.13421 [pdf, html, other]
Title: Detecting Machine-Generated Music with Explainability -- A Challenge and Early Benchmarks
Yupei Li, Qiyang Sun, Hanqian Li, Lucia Specia, Björn W. Schuller
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[84] arXiv:2412.13462 [pdf, html, other]
Title: SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
Kazuki Shimada, Christian Simon, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji
Comments: 5 pages, 3 figures
Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[85] arXiv:2412.13514 [pdf, html, other]
Title: Tuning Music Education: AI-Powered Personalization in Learning Music
Mayank Sanganeria, Rohan Gala
Comments: 38th Conference on Neural Information Processing Systems (NeurIPS 2024) Creative AI Track
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[86] arXiv:2412.15023 [pdf, html, other]
Title: FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Riccardo Fosco Gramaccioni, Christian Marinoni, Emilian Postolache, Marco Comunità, Luca Cosmo, Joshua D. Reiss, Danilo Comminiello
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[87] arXiv:2412.15230 [pdf, html, other]
Title: Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS Challenge
Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar, Sophie Young, Yao Xiao, Hend Elghazaly, Fritz Peters, Caitlin Illingworth, Dorota Braun, Ronan O'Malley, Simon Bell, Daniel Blackburn, Fasih Haider, Saturnino Luz, Heidi Christensen
Comments: 2 pages, no figure, conference
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[88] arXiv:2412.15602 [pdf, other]
Title: Music Genre Classification: Ensemble Learning with Subcomponents-level Attention
Yichen Liu, Abhijit Dasgupta, Qiwei He
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[89] arXiv:2412.16176 [pdf, html, other]
Title: Efficient VoIP Communications through LLM-based Real-Time Speech Reconstruction and Call Prioritization for Emergency Services
Danush Venkateshperumal, Rahman Abdul Rafi, Shakil Ahmed, Ashfaq Khokhar
Comments: 15 pages,8 figures
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[90] arXiv:2412.16182 [pdf, other]
Title: Decoding Poultry Vocalizations -- Natural Language Processing and Transformer Models for Semantic and Emotional Analysis
Venkatraman Manikandan, Suresh Neethirajan
Comments: 28 Pages, 14 figures
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[91] arXiv:2412.16267 [pdf, html, other]
Title: A Classification Benchmark for Artificial Intelligence Detection of Laryngeal Cancer from Patient Voice
Mary Paterson, James Moor, Luisa Cutillo
Comments: 16 pages, 6 figures, 10 tables
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Quantitative Methods (q-bio.QM)
[92] arXiv:2412.16526 [pdf, html, other]
Title: Text2midi: Generating Symbolic Music from Captions
Keshav Bhandari, Abhinaba Roy, Kyra Wang, Geeta Puri, Simon Colton, Dorien Herremans
Comments: 9 pages, 3 figures, Accepted at the 39th AAAI Conference on Artificial Intelligence (AAAI 2025)
Journal-ref: Proceedings of the 39th AAAI Conference on Artificial Intelligence (AAAI 2025)
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[93] arXiv:2412.16530 [pdf, html, other]
Title: Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
Lucas Goncalves, Prashant Mathur, Xing Niu, Brady Houston, Chandrashekhar Lavania, Srikanth Vishnubhotla, Lijia Sun, Anthony Ferritto
Comments: Accepted at ICASSP, 4 pages
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[94] arXiv:2412.16626 [pdf, html, other]
Title: Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
Junyu Wang, Zizhen Lin, Tianrui Wang, Meng Ge, Longbiao Wang, Jianwu Dang
Comments: Accepted at ICASSP 2025, 5 pages, 1 figures, 5 tables
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[95] arXiv:2412.16861 [pdf, html, other]
Title: SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
Yuhang He, Sangyun Shin, Anoop Cherian, Niki Trigoni, Andrew Markham
Comments: Accepted by WACV2025
Journal-ref: WACV2025
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[96] arXiv:2412.16904 [pdf, html, other]
Title: Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition
Jiaqi Zhao, Fei Wang, Kun Li, Yanyan Wei, Shengeng Tang, Shu Zhao, Xiao Sun
Comments: Accepted by ICASSP 2025
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[97] arXiv:2412.16928 [pdf, html, other]
Title: AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification
Zhenyuan Xiao, Yizhuo Yang, Guili Xu, Xianglong Zeng, Shenghai Yuan
Comments: Submitted to ICRA 2025
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[98] arXiv:2412.17212 [pdf, html, other]
Title: Trainingless Adaptation of Pretrained Models for Environmental Sound Classification
Noriyuki Tonami, Wataru Kohno, Keisuke Imoto, Yoshiyuki Yajima, Sakiko Mishima, Reishi Kondo, Tomoyuki Hino
Comments: Accepted to ICASSP2025
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[99] arXiv:2412.17306 [pdf, html, other]
Title: Multiple Consistency-guided Test-Time Adaptation for Contrastive Audio-Language Models with Unlabeled Audio
Gongyu Chen, Haomin Zhang, Chaofan Ding, Zihao Chen, Xinhan Di
Comments: 6 pages, 1 figure, accepted by ICASSP 2025
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[100] arXiv:2412.17667 [pdf, html, other]
Title: VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
Jiatong Shi, Hye-jin Shim, Jinchuan Tian, Siddhant Arora, Haibin Wu, Darius Petermann, Jia Qi Yip, You Zhang, Yuxun Tang, Wangyou Zhang, Dareen Safar Alharthi, Yichen Huang, Koichi Saito, Jionghao Han, Yiwen Zhao, Chris Donahue, Shinji Watanabe
Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[101] arXiv:2412.17924 [pdf, html, other]
Title: Are audio DeepFake detection models polyglots?
Bartłomiej Marek, Piotr Kawa, Piotr Syga
Comments: Keywords: Audio DeepFakes, DeepFake detection, multilingual audio DeepFakes
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[102] arXiv:2412.18061 [pdf, html, other]
Title: Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
Hyunbae Jeon, Frederic Guintu, Rayvant Sahni
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Audio and Speech Processing (eess.AS)
[103] arXiv:2412.18157 [pdf, html, other]
Title: Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
Yaoyun Zhang, Xuenan Xu, Mengyue Wu
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[104] arXiv:2412.18191 [pdf, html, other]
Title: Explaining Speaker and Spoof Embeddings via Probing
Xuechen Liu, Junichi Yamagishi, Md Sahidullah, Tomi kinnunen
Comments: To appear in IEEE ICASSP 2025
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[105] arXiv:2412.18217 [pdf, html, other]
Title: U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation
Shaoxiang Dang, Tetsuya Matsumoto, Yoshinori Takeuchi, Hiroaki Kudo
Journal-ref: 2024 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[106] arXiv:2412.18710 [pdf, other]
Title: Simi-SFX: A similarity-based conditioning method for controllable sound effect synthesis
Yunyi Liu, Craig Jin
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[107] arXiv:2412.18836 [pdf, html, other]
Title: MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
Neil Shah, Ayan Kashyap, Shirish Karande, Vineet Gandhi
Comments: Accepted at IEEE ICASSP 2025
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[108] arXiv:2412.18839 [pdf, html, other]
Title: Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset
Neil Shah, Shirish Karande, Vineet Gandhi
Comments: Accepted at IEEE ICASSP 2025
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[109] arXiv:2412.18851 [pdf, html, other]
Title: Attention-Enhanced Short-Time Wiener Solution for Acoustic Echo Cancellation
Fei Zhao, Xueliang Zhang
Subjects: Sound (cs.SD)
[110] arXiv:2412.18913 [pdf, html, other]
Title: Robust Target Speaker Direction of Arrival Estimation
Zixuan Li, Shulin He, Xueliang Zhang
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[111] arXiv:2412.18955 [pdf, html, other]
Title: Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations
Julien Guinot, Elio Quinton, György Fazekas
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[112] arXiv:2412.19099 [pdf, html, other]
Title: BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement
Cunhang Fan, Enrui Liu, Andong Li, Jianhua Tao, Jian Zhou, Jiahao Li, Chengshi Zheng, Zhao Lv
Comments: Accepted by AAAI 2025
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[113] arXiv:2412.19123 [pdf, html, other]
Title: CoheDancers: Enhancing Interactive Group Dance Generation through Music-Driven Coherence Decomposition
Kaixing Yang, Xulong Tang, Haoyu Wu, Qinliang Xue, Biao Qin, Hongyan Liu, Zhaoxin Fan
Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[114] arXiv:2412.19200 [pdf, html, other]
Title: Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
Dengming Zhang, Weitao You, Ziheng Liu, Lingyun Sun, Pei Chen
Comments: Accepted by the 39th AAAI Conference on Artificial Intelligence (AAAI-25)
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[115] arXiv:2412.19279 [pdf, html, other]
Title: Improving Generalization for AI-Synthesized Voice Detection
Hainan Ren, Li Lin, Chun-Hao Liu, Xin Wang, Shu Hu
Comments: AAAI25
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[116] arXiv:2412.19351 [pdf, other]
Title: ETTA: Elucidating the Design Space of Text-to-Audio Models
Sang-gil Lee, Zhifeng Kong, Arushi Goel, Sungwon Kim, Rafael Valle, Bryan Catanzaro
Comments: ICML 2025. Demo: this https URL Code: this https URL
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[117] arXiv:2412.19909 [pdf, html, other]
Title: Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
Shreya G. Upadhyay, Ali N. Salman, Carlos Busso, Chi-Chun Lee
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[118] arXiv:2412.20155 [pdf, html, other]
Title: Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
Wooseok Han, Minki Kang, Changhun Kim, Eunho Yang
Comments: Accepted by ICASSP 2025
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[119] arXiv:2412.20914 [pdf, html, other]
Title: Language-based Audio Retrieval with Co-Attention Networks
Haoran Sun, Zimu Wang, Qiuyi Chen, Jianjun Chen, Jia Wang, Haiyang Zhang
Comments: Accepted at UIC 2024 proceedings. Accepted version
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[120] arXiv:2412.21037 [pdf, html, other]
Title: TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
Chia-Yu Hung, Navonil Majumder, Zhifeng Kong, Ambuj Mehrish, Amir Ali Bagherzadeh, Chuan Li, Rafael Valle, Bryan Catanzaro, Soujanya Poria
Comments: this https URL
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[121] arXiv:2412.00049 (cross-list from cs.MM) [pdf, html, other]
Title: A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning
Luis Vilaca, Yi Yu, Paula Vinan
Comments: arXiv admin note: text overlap with arXiv:2202.13673
Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[122] arXiv:2412.00055 (cross-list from eess.AS) [pdf, html, other]
Title: High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
Sourav Banerjee, Ayushi Agarwal, Promila Ghosh
Comments: 15 pages
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[123] arXiv:2412.00057 (cross-list from eess.AS) [pdf, html, other]
Title: Feasibility of Mental Health Triage Call Priority Prediction Using Machine Learning
Rajib Rana, Niall Higgins, Kazi Nazmul Haque, John Reilly, Kylie Burke, Kathryn Turner, Terry Stedman
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[124] arXiv:2412.00175 (cross-list from cs.CV) [pdf, html, other]
Title: Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
Stefan Smeu, Dragos-Alexandru Boldisor, Dan Oneata, Elisabeta Oneata
Comments: Accepted as a highlight paper at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[125] arXiv:2412.00721 (cross-list from cs.AI) [pdf, other]
Title: A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
Zheshu Song, Ziyang Ma, Yifan Yang, Jianheng Zhuo, Xie Chen
Comments: This work hasn't been finished yet
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Total of 231 entries : 26-125 101-200 201-231
Showing up to 100 entries per page: fewer | more | all
  • About
  • Help
  • contact arXivClick here to contact arXiv Contact
  • subscribe to arXiv mailingsClick here to subscribe Subscribe
  • Copyright
  • Privacy Policy
  • Web Accessibility Assistance
  • arXiv Operational Status
    Get status notifications via email or slack