Audio and Speech Processing

Authors and titles for March 2024

Total of 213 entries : 1-50 51-100 101-150 151-200 201-213

Showing up to 50 entries per page: fewer | more | all

[101] arXiv:2403.03762 (cross-list from eess.SP) [pdf, html, other]: Title: Room Impulse Response Estimation using Optimal Transport: Simulation-Informed Inference

David Sundström, Anton Björkman, Andreas Jakobsson, Filip Elvander

Subjects: Signal Processing (eess.SP); Audio and Speech Processing (eess.AS)
[102] arXiv:2403.03947 (cross-list from cs.SD) [pdf, html, other]: Title: Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset

Pedro Ramoneda, Minhee Lee, Dasaem Jeong, J.J. Valero-Mas, Xavier Serra

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[103] arXiv:2403.04111 (cross-list from cs.SD) [pdf, other]: Title: Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication

Yejin Jeon, Gary Geunbae Lee

Comments: Accepted to EACL Main 2024

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[104] arXiv:2403.04178 (cross-list from cs.CL) [pdf, html, other]: Title: Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation

Sai Akarsh, Vamshi Raghusimha, Anindita Mondal, Anil Vuppala

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[105] arXiv:2403.04245 (cross-list from cs.SD) [pdf, html, other]: Title: A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition

Yusheng Dai, Hang Chen, Jun Du, Ruoyu Wang, Shihao Chen, Jiefeng Ma, Haotian Wang, Chin-Hui Lee

Comments: the paper is accepted by CVPR2024

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[106] arXiv:2403.04594 (cross-list from cs.SD) [pdf, html, other]: Title: A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds

Xuenan Xu, Xiaohang Xu, Zeyu Xie, Pingyue Zhang, Mengyue Wu, Kai Yu

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[107] arXiv:2403.04654 (cross-list from cs.CV) [pdf, html, other]: Title: Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention

R. Gnana Praveen, Jahangir Alam

Comments: Accepted to FG2024

Subjects: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[108] arXiv:2403.04661 (cross-list from cs.CV) [pdf, html, other]: Title: Dynamic Cross Attention for Audio-Visual Person Verification

R. Gnana Praveen, Jahangir Alam

Comments: Accepted to FG2024

Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[109] arXiv:2403.05010 (cross-list from cs.SD) [pdf, html, other]: Title: RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction

Peng Liu, Dongyang Dai, Zhiyong Wu

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[110] arXiv:2403.05380 (cross-list from cs.SD) [pdf, html, other]: Title: Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings

Mahyar Gohari, Paolo Bestagini, Sergio Benini, Nicola Adami

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[111] arXiv:2403.05583 (cross-list from cs.HC) [pdf, html, other]: Title: A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition

Tyler Benster, Guy Wilson, Reshef Elisha, Francis R Willett, Shaul Druckmann

Subjects: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[112] arXiv:2403.05772 (cross-list from cs.SD) [pdf, html, other]: Title: sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks

Qu Yang, Qianhui Liu, Nan Li, Meng Ge, Zeyang Song, Haizhou Li

Comments: Accepted by ICASSP 2024

Subjects: Sound (cs.SD); Neural and Evolutionary Computing (cs.NE); Audio and Speech Processing (eess.AS)
[113] arXiv:2403.05820 (cross-list from cs.SD) [pdf, html, other]: Title: An Audio-textual Diffusion Model For Converting Speech Signals Into Ultrasound Tongue Imaging Data

Yudong Yang, Rongfeng Su, Xiaokang Liu, Nan Yan, Lan Wang

Comments: ICASSP2024 Accept

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[114] arXiv:2403.05834 (cross-list from cs.MM) [pdf, html, other]: Title: Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information

Qiaochu Huang, Xu He, Boshi Tang, Haolin Zhuang, Liyang Chen, Shuochen Gao, Zhiyong Wu, Haozhi Huang, Helen Meng

Subjects: Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[115] arXiv:2403.05989 (cross-list from cs.SD) [pdf, html, other]: Title: HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling

Chunhui Wang, Chang Zeng, Bowen Zhang, Ziyang Ma, Yefan Zhu, Zifeng Cai, Jian Zhao, Zhonglin Jiang, Yong Chen

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[116] arXiv:2403.06100 (cross-list from cs.HC) [pdf, html, other]: Title: Automatic design optimization of preference-based subjective evaluation with online learning in crowdsourcing environment

Yusuke Yasuda, Tomoki Toda

Subjects: Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
[117] arXiv:2403.06260 (cross-list from cs.CL) [pdf, html, other]: Title: SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations

Amit Meghanani, Thomas Hain

Comments: Accepted at ICASSP 2024

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[118] arXiv:2403.06387 (cross-list from cs.SD) [pdf, html, other]: Title: Towards Decoupling Frontend Enhancement and Backend Recognition in Monaural Robust ASR

Yufeng Yang, Ashutosh Pandey, DeLiang Wang

Comments: Submitted to IEEE/ACM Transactions on Audio, Speech and Language Processing. arXiv admin note: text overlap with arXiv:2210.13318

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[119] arXiv:2403.06404 (cross-list from cs.SD) [pdf, html, other]: Title: Cosine Scoring with Uncertainty for Neural Speaker Embedding

Qiongqiong Wang, Kong Aik Lee

Comments: 5 pages, 4 figures

Journal-ref: IEEE Signal Processing Letters 2024

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[120] arXiv:2403.06487 (cross-list from cs.CL) [pdf, html, other]: Title: Multilingual Turn-taking Prediction Using Voice Activity Projection

Koji Inoue, Bing'er Jiang, Erik Ekstedt, Tatsuya Kawahara, Gabriel Skantze

Comments: This paper has been accepted for presentation at The 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) and represents the author's version of the work

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[121] arXiv:2403.07675 (cross-list from cs.SD) [pdf, html, other]: Title: Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers

Changsheng Quan, Xiaofei Li

Comments: Accepted by IEEE Signal Processing Letters

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[122] arXiv:2403.07802 (cross-list from cs.SD) [pdf, html, other]: Title: Boosting keyword spotting through on-device learnable user speech characteristics

Cristian Cioflan, Lukas Cavigelli, Luca Benini

Comments: 5 pages, 3 tables, 2 figures. Accepted as a full paper by the tinyML Research Symposium 2024

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[123] arXiv:2403.07938 (cross-list from cs.SD) [pdf, html, other]: Title: Text-to-Audio Generation Synchronized with Videos

Shentong Mo, Jing Shi, Yapeng Tian

Comments: arXiv admin note: text overlap with arXiv:2305.12903

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[124] arXiv:2403.07995 (cross-list from cs.SD) [pdf, html, other]: Title: Motifs, Phrases, and Beyond: The Modelling of Structure in Symbolic Music Generation

Keshav Bhandari, Simon Colton

Comments: Accepted to 13th International Conference on Artificial Intelligence in Music, Sound, Art and Design (EvoMUSART) 2024

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Symbolic Computation (cs.SC); Audio and Speech Processing (eess.AS)
[125] arXiv:2403.08164 (cross-list from cs.SD) [pdf, html, other]: Title: EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech

Ziqi Liang, Haoxiang Shi, Jiawei Wang, Keda Lu

Comments: Accepted by the 27th IEEE International Conference on Computer Supported Cooperative Work in Design (IEEE CSCWD 2024). arXiv admin note: substantial text overlap with arXiv:2211.01948

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[126] arXiv:2403.08187 (cross-list from cs.CL) [pdf, html, other]: Title: Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children

Taekyung Ahn, Yeonjung Hong, Younggon Im, Do Hyung Kim, Dayoung Kang, Joo Won Jeong, Jae Won Kim, Min Jung Kim, Ah-ra Cho, Dae-Hyun Jang, Hosung Nam

Comments: 12 pages, 2 figures

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[127] arXiv:2403.08196 (cross-list from cs.CL) [pdf, html, other]: Title: SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation

Jiayu Du, Jinpeng Li, Guoguo Chen, Wei-Qiang Zhang

Journal-ref: Computer Speech & Language (2025)

Subjects: Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[128] arXiv:2403.08525 (cross-list from cs.SD) [pdf, html, other]: Title: From Weak to Strong Sound Event Labels using Adaptive Change-Point Detection and Active Learning

John Martinsson, Olof Mogren, Maria Sandsten, Tuomas Virtanen

Comments: Accepted at EUSIPCO 2024 (nominated best student paper)

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[129] arXiv:2403.08559 (cross-list from cs.SD) [pdf, html, other]: Title: End-to-End Amp Modeling: From Data to Controllable Guitar Amplifier Models

Lauri Juvela, Eero-Pekka Damskägg, Aleksi Peussa, Jaakko Mäkinen, Thomas Sherson, Stylianos I. Mimilakis, Athanasios Gotsopoulos

Comments: Presented at ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[130] arXiv:2403.08738 (cross-list from cs.CL) [pdf, html, other]: Title: Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations

Amit Meghanani, Thomas Hain

Comments: Accepted to EACL 2024 Main Conference, Long paper

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[131] arXiv:2403.09030 (cross-list from cs.SD) [pdf, other]: Title: An AI-Driven Approach to Wind Turbine Bearing Fault Diagnosis from Acoustic Signals

Zhao Wang, Xiaomeng Li, Na Li, Longlong Shu

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[132] arXiv:2403.09298 (cross-list from cs.SD) [pdf, html, other]: Title: More than words: Advancements and challenges in speech recognition for singing

Anna Kruspe

Comments: Conference on Electronic Speech Signal Processing (ESSV) 2024, Keynote

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[133] arXiv:2403.09321 (cross-list from cs.SD) [pdf, html, other]: Title: A Practical Guide to Spectrogram Analysis for Audio Signal Processing

Zulfidin Khodzhaev

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[134] arXiv:2403.09407 (cross-list from cs.SD) [pdf, html, other]: Title: LM2D: Lyrics- and Music-Driven Dance Synthesis

Wenjie Yin, Xuejiao Zhao, Yi Yu, Hang Yin, Danica Kragic, Mårten Björkman

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[135] arXiv:2403.09451 (cross-list from cs.CV) [pdf, html, other]: Title: M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment

Long Nguyen-Phuoc, Renald Gaboriau, Dimitri Delacroix, Laurent Navarro

Journal-ref: Proceedings of the 19th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - Volume 2 VISAPP: VISAPP, 869-876, 2024 , Rome, Italy

Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[136] arXiv:2403.09455 (cross-list from cs.SD) [pdf, html, other]: Title: The Neural-SRP method for positional sound source localization

Eric Grinstein, Toon van Waterschoot, Mike Brookes, Patrick A. Naylor

Comments: Presented at Asilomar Conference on Signals, Systems, and Computers

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[137] arXiv:2403.09579 (cross-list from cs.SD) [pdf, html, other]: Title: uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures

Afrina Tabassum, Dung Tran, Trung Dang, Ismini Lourentzou, Kazuhito Koishida

Comments: 5 pages, 6 figures, 4 tables. To appear in ICASSP'2024

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[138] arXiv:2403.09598 (cross-list from cs.SD) [pdf, html, other]: Title: Mixture of Mixups for Multi-label Classification of Rare Anuran Sounds

Ilyass Moummad, Nicolas Farrugia, Romain Serizel, Jeremy Froidevaux, Vincent Lostanlen

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[139] arXiv:2403.09753 (cross-list from cs.SD) [pdf, html, other]: Title: SpokeN-100: A Cross-Lingual Benchmarking Dataset for The Classification of Spoken Numbers in Different Languages

René Groh, Nina Goes, Andreas M. Kist

Comments: Accepted as a full paper by the tinyML Research Symposium 2024

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[140] arXiv:2403.10024 (cross-list from cs.SD) [pdf, html, other]: Title: MR-MT3: Memory Retaining Multi-Track Music Transcription to Mitigate Instrument Leakage

Hao Hao Tan, Kin Wai Cheuk, Taemin Cho, Wei-Hsiang Liao, Yuki Mitsufuji

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[141] arXiv:2403.10146 (cross-list from cs.SD) [pdf, html, other]: Title: Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval

Qian Wang, Jia-Chen Gu, Zhen-Hua Ling

Comments: 5 pages, accepted to ICASSP2024

Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[142] arXiv:2403.10329 (cross-list from eess.SP) [pdf, html, other]: Title: Multi-Source Localization and Data Association for Time-Difference of Arrival Measurements

Gabrielle Flood, Filip Elvander

Subjects: Signal Processing (eess.SP); Sound (cs.SD); Audio and Speech Processing (eess.AS); Optimization and Control (math.OC)
[143] arXiv:2403.10380 (cross-list from cs.SD) [pdf, other]: Title: BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics

Lukas Rauch, Raphael Schwinger, Moritz Wirth, René Heinrich, Denis Huseljic, Marek Herde, Jonas Lange, Stefan Kahl, Bernhard Sick, Sven Tomforde, Christoph Scholz

Comments: accepted as spotlight @ICLR2025

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[144] arXiv:2403.10488 (cross-list from cs.CV) [pdf, html, other]: Title: Joint Multimodal Transformer for Emotion Recognition in the Wild

Paul Waligora, Haseeb Aslam, Osama Zeeshan, Soufiane Belharbi, Alessandro Lameiras Koerich, Marco Pedersoli, Simon Bacon, Eric Granger

Comments: 10 pages, 4 figures, 6 tables, CVPRw 2024

Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[145] arXiv:2403.10493 (cross-list from cs.SD) [pdf, html, other]: Title: MusicHiFi: Fast High-Fidelity Stereo Vocoding

Ge Zhu, Juan-Pablo Caceres, Zhiyao Duan, Nicholas J. Bryan

Comments: Accepted to IEEE Signal Processing Letters

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
[146] arXiv:2403.10518 (cross-list from cs.CV) [pdf, html, other]: Title: Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives

Ronghui Li, YuXiang Zhang, Yachao Zhang, Hongwen Zhang, Jie Guo, Yan Zhang, Yebin Liu, Xiu Li

Comments: Accepted by CVPR2024, Project page: this https URL

Subjects: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[147] arXiv:2403.10549 (cross-list from cs.SD) [pdf, html, other]: Title: On-Device Domain Learning for Keyword Spotting on Low-Power Extreme Edge Embedded Systems

Cristian Cioflan, Lukas Cavigelli, Manuele Rusci, Miguel de Prado, Luca Benini

Comments: 5 pages, 2 tables, 2 figures. Accepted at IEEE AICAS 2024

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[148] arXiv:2403.10796 (cross-list from cs.SD) [pdf, html, other]: Title: CoPlay: Audio-agnostic Cognitive Scaling for Acoustic Sensing

Yin Li, Rajalakshmi Nanadakumar

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[149] arXiv:2403.10805 (cross-list from cs.SD) [pdf, other]: Title: Speech-driven Personalized Gesture Synthetics: Harnessing Automatic Fuzzy Feature Inference

Fan Zhang, Zhaohan Wang, Xin Lyu, Siyuan Zhao, Mengjian Li, Weidong Geng, Naye Ji, Hui Du, Fuxing Gao, Hao Wu, Shunman Li

Comments: 12 pages,

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Human-Computer Interaction (cs.HC); Audio and Speech Processing (eess.AS)
[150] arXiv:2403.10904 (cross-list from cs.SD) [pdf, html, other]: Title: Urban Sound Propagation: a Benchmark for 1-Step Generative Modeling of Complex Physical Systems

Martin Spitznagel, Janis Keuper

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)

Total of 213 entries : 1-50 51-100 101-150 151-200 201-213

Showing up to 50 entries per page: fewer | more | all