close this message
arXiv smileybones

Happy Birthday to arXiv!

It's our birthday — woohoo! On August 14th, 1991, the very first paper was submitted to arXiv. That's 34 years of open science! Give today and help support arXiv for many birthdays to come.

Give a gift!
Skip to main content
Cornell University
We gratefully acknowledge support from the Simons Foundation, member institutions, and all contributors. Donate
arxiv logo > eess.AS

Help | Advanced Search

arXiv logo
Cornell University Logo

quick links

  • Login
  • Help Pages
  • About

Audio and Speech Processing

Authors and titles for March 2024

Total of 213 entries : 1-25 26-50 51-75 76-100 101-125 126-150 151-175 176-200 ... 201-213
Showing up to 25 entries per page: fewer | more | all
[101] arXiv:2403.03762 (cross-list from eess.SP) [pdf, html, other]
Title: Room Impulse Response Estimation using Optimal Transport: Simulation-Informed Inference
David Sundström, Anton Björkman, Andreas Jakobsson, Filip Elvander
Subjects: Signal Processing (eess.SP); Audio and Speech Processing (eess.AS)
[102] arXiv:2403.03947 (cross-list from cs.SD) [pdf, html, other]
Title: Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset
Pedro Ramoneda, Minhee Lee, Dasaem Jeong, J.J. Valero-Mas, Xavier Serra
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[103] arXiv:2403.04111 (cross-list from cs.SD) [pdf, other]
Title: Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication
Yejin Jeon, Gary Geunbae Lee
Comments: Accepted to EACL Main 2024
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[104] arXiv:2403.04178 (cross-list from cs.CL) [pdf, html, other]
Title: Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
Sai Akarsh, Vamshi Raghusimha, Anindita Mondal, Anil Vuppala
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[105] arXiv:2403.04245 (cross-list from cs.SD) [pdf, html, other]
Title: A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition
Yusheng Dai, Hang Chen, Jun Du, Ruoyu Wang, Shihao Chen, Jiefeng Ma, Haotian Wang, Chin-Hui Lee
Comments: the paper is accepted by CVPR2024
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[106] arXiv:2403.04594 (cross-list from cs.SD) [pdf, html, other]
Title: A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
Xuenan Xu, Xiaohang Xu, Zeyu Xie, Pingyue Zhang, Mengyue Wu, Kai Yu
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[107] arXiv:2403.04654 (cross-list from cs.CV) [pdf, html, other]
Title: Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention
R. Gnana Praveen, Jahangir Alam
Comments: Accepted to FG2024
Subjects: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[108] arXiv:2403.04661 (cross-list from cs.CV) [pdf, html, other]
Title: Dynamic Cross Attention for Audio-Visual Person Verification
R. Gnana Praveen, Jahangir Alam
Comments: Accepted to FG2024
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[109] arXiv:2403.05010 (cross-list from cs.SD) [pdf, html, other]
Title: RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction
Peng Liu, Dongyang Dai, Zhiyong Wu
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[110] arXiv:2403.05380 (cross-list from cs.SD) [pdf, html, other]
Title: Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings
Mahyar Gohari, Paolo Bestagini, Sergio Benini, Nicola Adami
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[111] arXiv:2403.05583 (cross-list from cs.HC) [pdf, html, other]
Title: A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
Tyler Benster, Guy Wilson, Reshef Elisha, Francis R Willett, Shaul Druckmann
Subjects: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[112] arXiv:2403.05772 (cross-list from cs.SD) [pdf, html, other]
Title: sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks
Qu Yang, Qianhui Liu, Nan Li, Meng Ge, Zeyang Song, Haizhou Li
Comments: Accepted by ICASSP 2024
Subjects: Sound (cs.SD); Neural and Evolutionary Computing (cs.NE); Audio and Speech Processing (eess.AS)
[113] arXiv:2403.05820 (cross-list from cs.SD) [pdf, html, other]
Title: An Audio-textual Diffusion Model For Converting Speech Signals Into Ultrasound Tongue Imaging Data
Yudong Yang, Rongfeng Su, Xiaokang Liu, Nan Yan, Lan Wang
Comments: ICASSP2024 Accept
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[114] arXiv:2403.05834 (cross-list from cs.MM) [pdf, html, other]
Title: Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information
Qiaochu Huang, Xu He, Boshi Tang, Haolin Zhuang, Liyang Chen, Shuochen Gao, Zhiyong Wu, Haozhi Huang, Helen Meng
Subjects: Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[115] arXiv:2403.05989 (cross-list from cs.SD) [pdf, html, other]
Title: HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
Chunhui Wang, Chang Zeng, Bowen Zhang, Ziyang Ma, Yefan Zhu, Zifeng Cai, Jian Zhao, Zhonglin Jiang, Yong Chen
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[116] arXiv:2403.06100 (cross-list from cs.HC) [pdf, html, other]
Title: Automatic design optimization of preference-based subjective evaluation with online learning in crowdsourcing environment
Yusuke Yasuda, Tomoki Toda
Subjects: Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
[117] arXiv:2403.06260 (cross-list from cs.CL) [pdf, html, other]
Title: SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations
Amit Meghanani, Thomas Hain
Comments: Accepted at ICASSP 2024
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[118] arXiv:2403.06387 (cross-list from cs.SD) [pdf, html, other]
Title: Towards Decoupling Frontend Enhancement and Backend Recognition in Monaural Robust ASR
Yufeng Yang, Ashutosh Pandey, DeLiang Wang
Comments: Submitted to IEEE/ACM Transactions on Audio, Speech and Language Processing. arXiv admin note: text overlap with arXiv:2210.13318
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[119] arXiv:2403.06404 (cross-list from cs.SD) [pdf, html, other]
Title: Cosine Scoring with Uncertainty for Neural Speaker Embedding
Qiongqiong Wang, Kong Aik Lee
Comments: 5 pages, 4 figures
Journal-ref: IEEE Signal Processing Letters 2024
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[120] arXiv:2403.06487 (cross-list from cs.CL) [pdf, html, other]
Title: Multilingual Turn-taking Prediction Using Voice Activity Projection
Koji Inoue, Bing'er Jiang, Erik Ekstedt, Tatsuya Kawahara, Gabriel Skantze
Comments: This paper has been accepted for presentation at The 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) and represents the author's version of the work
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[121] arXiv:2403.07675 (cross-list from cs.SD) [pdf, html, other]
Title: Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers
Changsheng Quan, Xiaofei Li
Comments: Accepted by IEEE Signal Processing Letters
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[122] arXiv:2403.07802 (cross-list from cs.SD) [pdf, html, other]
Title: Boosting keyword spotting through on-device learnable user speech characteristics
Cristian Cioflan, Lukas Cavigelli, Luca Benini
Comments: 5 pages, 3 tables, 2 figures. Accepted as a full paper by the tinyML Research Symposium 2024
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[123] arXiv:2403.07938 (cross-list from cs.SD) [pdf, html, other]
Title: Text-to-Audio Generation Synchronized with Videos
Shentong Mo, Jing Shi, Yapeng Tian
Comments: arXiv admin note: text overlap with arXiv:2305.12903
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[124] arXiv:2403.07995 (cross-list from cs.SD) [pdf, html, other]
Title: Motifs, Phrases, and Beyond: The Modelling of Structure in Symbolic Music Generation
Keshav Bhandari, Simon Colton
Comments: Accepted to 13th International Conference on Artificial Intelligence in Music, Sound, Art and Design (EvoMUSART) 2024
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Symbolic Computation (cs.SC); Audio and Speech Processing (eess.AS)
[125] arXiv:2403.08164 (cross-list from cs.SD) [pdf, html, other]
Title: EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech
Ziqi Liang, Haoxiang Shi, Jiawei Wang, Keda Lu
Comments: Accepted by the 27th IEEE International Conference on Computer Supported Cooperative Work in Design (IEEE CSCWD 2024). arXiv admin note: substantial text overlap with arXiv:2211.01948
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
Total of 213 entries : 1-25 26-50 51-75 76-100 101-125 126-150 151-175 176-200 ... 201-213
Showing up to 25 entries per page: fewer | more | all
  • About
  • Help
  • contact arXivClick here to contact arXiv Contact
  • subscribe to arXiv mailingsClick here to subscribe Subscribe
  • Copyright
  • Privacy Policy
  • Web Accessibility Assistance
  • arXiv Operational Status
    Get status notifications via email or slack