Skip to main content
Cornell University
We gratefully acknowledge support from the Simons Foundation, member institutions, and all contributors. Donate
arxiv logo > cs.SD

Help | Advanced Search

arXiv logo
Cornell University Logo

quick links

  • Login
  • Help Pages
  • About

Sound

Authors and titles for September 2023

Total of 451 entries : 1-50 51-100 101-150 151-200 201-250 251-300 301-350 ... 451-451
Showing up to 50 entries per page: fewer | more | all
[151] arXiv:2309.14158 [pdf, other]
Title: An Investigation of Distribution Alignment in Multi-Genre Speaker Recognition
Zhenyu Zhou, Junhui Chen, Namin Wang, Lantian Li, Dong Wang
Comments: submitted to ICASSP 2024
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[152] arXiv:2309.14383 [pdf, other]
Title: Towards using Cough for Respiratory Disease Diagnosis by leveraging Artificial Intelligence: A Survey
Aneeqa Ijaz, Muhammad Nabeel, Usama Masood, Tahir Mahmood, Mydah Sajid Hashmi, Iryna Posokhova, Ali Rizwan, Ali Imran
Comments: 30 pages, 12 figures, 9 tables
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[153] arXiv:2309.14405 [pdf, html, other]
Title: Joint Audio and Speech Understanding
Yuan Gong, Alexander H. Liu, Hongyin Luo, Leonid Karlinsky, James Glass
Comments: Accepted at ASRU 2023. Code, dataset, and pretrained models are at this https URL. Interactive demo at this https URL
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[154] arXiv:2309.14586 [pdf, other]
Title: Speech Audio Synthesis from Tagged MRI and Non-Negative Matrix Factorization via Plastic Transformer
Xiaofeng Liu, Fangxu Xing, Maureen Stone, Jiachen Zhuo, Sidney Fels, Jerry L. Prince, Georges El Fakhri, Jonghye Woo
Comments: MICCAI 2023 (Oral presentation)
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
[155] arXiv:2309.14838 [pdf, html, other]
Title: Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification
Duc-Tuan Truong, Ruijie Tao, Jia Qi Yip, Kong Aik Lee, Eng Siong Chng
Comments: Accepted by ICASSP 2024
Journal-ref: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 10336-10340
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[156] arXiv:2309.15024 [pdf, other]
Title: Synthia's Melody: A Benchmark Framework for Unsupervised Domain Adaptation in Audio
Chia-Hsin Lin, Charles Jones, Björn W. Schuller, Harry Coppock
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[157] arXiv:2309.15512 [pdf, html, other]
Title: High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models
Chunyu Qiang, Hao Li, Yixin Tian, Yi Zhao, Ying Zhang, Longbiao Wang, Jianwu Dang
Comments: Accepted by ICASSP 2024. arXiv admin note: substantial text overlap with arXiv:2307.15484; text overlap with arXiv:2309.00424
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[158] arXiv:2309.15674 [pdf, other]
Title: Speech collage: code-switched audio generation by collaging monolingual corpora
Amir Hussein, Dorsa Zeinali, Ondřej Klejch, Matthew Wiesner, Brian Yan, Shammur Chowdhury, Ahmed Ali, Shinji Watanabe, Sanjeev Khudanpur
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[159] arXiv:2309.15977 [pdf, other]
Title: Neural Acoustic Context Field: Rendering Realistic Room Impulse Response With Neural Fields
Susan Liang, Chao Huang, Yapeng Tian, Anurag Kumar, Chenliang Xu
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[160] arXiv:2309.16178 [pdf, other]
Title: LAE-ST-MoE: Boosted Language-Aware Encoder Using Speech Translation Auxiliary Task for E2E Code-switching ASR
Guodong Ma, Wenxuan Wang, Yuke Li, Yuting Yang, Binbin Du, Haoran Fu
Comments: Accepted to IEEE ASRU 2023
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[161] arXiv:2309.16265 [pdf, other]
Title: Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
Wuyang Liu, Yanzhen Ren
Comments: 5 pages, 3 figures. Accepted by ICASSP 2024
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[162] arXiv:2309.16284 [pdf, html, other]
Title: NOMAD: Unsupervised Learning of Perceptual Embeddings for Speech Enhancement and Non-matching Reference Audio Quality Assessment
Alessandro Ragano, Jan Skoglund, Andrew Hines
Comments: Accepted for ICASSP 2024
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[163] arXiv:2309.16287 [pdf, other]
Title: Predicting performance difficulty from piano sheet music images
Pedro Ramoneda, Jose J. Valero-Mas, Dasaem Jeong, Xavier Serra
Subjects: Sound (cs.SD); Digital Libraries (cs.DL); Audio and Speech Processing (eess.AS)
[164] arXiv:2309.16369 [pdf, other]
Title: Bringing the Discussion of Minima Sharpness to the Audio Domain: a Filter-Normalised Evaluation for Acoustic Scene Classification
Manuel Milling, Andreas Triantafyllopoulos, Iosif Tsangko, Simon David Noel Rampp, Björn Wolfgang Schuller
Comments: This work has been submitted to the IEEE for possible publication
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[165] arXiv:2309.16418 [pdf, other]
Title: Efficient Supervised Training of Audio Transformers for Music Representation Learning
Pablo Alonso-Jiménez, Xavier Serra, Dmitry Bogdanov
Comments: Accepted at the 2023 International Society for Music Information Retrieval Conference (ISMIR'23)
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[166] arXiv:2309.16569 [pdf, other]
Title: Audio-Visual Speaker Verification via Joint Cross-Attention
R. Gnana Praveen, Jahangir Alam
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[167] arXiv:2309.17056 [pdf, html, other]
Title: ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
Wenhao Guan, Qi Su, Haodong Zhou, Shiyu Miao, Xingjia Xie, Lin Li, Qingyang Hong
Comments: Accepted at ICASSP2024
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[168] arXiv:2309.17189 [pdf, html, other]
Title: RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
Samuel Pegg, Kai Li, Xiaolin Hu
Comments: Accepted by The Twelfth International Conference on Learning Representations (ICLR) 2024, see this https URL
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[169] arXiv:2309.17352 [pdf, html, other]
Title: Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
Shih-Lun Wu, Xuankai Chang, Gordon Wichern, Jee-weon Jung, François Germain, Jonathan Le Roux, Shinji Watanabe
Comments: ICASSP 2024 camera-ready paper. Winner of the DCASE 2023 Challenge Task 6A: Automated Audio Captioning (AAC)
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[170] arXiv:2309.00169 (cross-list from eess.AS) [pdf, html, other]
Title: RepCodec: A Speech Representation Codec for Speech Tokenization
Zhichao Huang, Chutong Meng, Tom Ko
Comments: ACL 2024 (Main)
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[171] arXiv:2309.00223 (cross-list from eess.AS) [pdf, html, other]
Title: The FruitShell French synthesis system at the Blizzard 2023 Challenge
Xin Qi, Xiaopeng Wang, Zhiyong Wang, Wang Liu, Mingming Ding, Shuchen Shi
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[172] arXiv:2309.00347 (cross-list from cs.IR) [pdf, other]
Title: Towards Contrastive Learning in Music Video Domain
Karel Veldkamp, Mariya Hendriksen, Zoltán Szlávik, Alexander Keijser
Comments: 6 pages, 2 figures, 2 tables
Subjects: Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[173] arXiv:2309.00376 (cross-list from eess.AS) [pdf, other]
Title: Remixing-based Unsupervised Source Separation from Scratch
Kohei Saijo, Tetsuji Ogawa
Comments: Interspeech2023, 5pages, 2figures, 2tables
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[174] arXiv:2309.00424 (cross-list from eess.AS) [pdf, html, other]
Title: Learning Speech Representation From Contrastive Token-Acoustic Pretraining
Chunyu Qiang, Hao Li, Yixin Tian, Ruibo Fu, Tao Wang, Longbiao Wang, Jianwu Dang
Comments: Accepted by ICASSP 2024
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)
[175] arXiv:2309.00647 (cross-list from eess.AS) [pdf, other]
Title: Improving Small Footprint Few-shot Keyword Spotting with Supervision on Auxiliary Data
Seunghan Yang, Byeonggeun Kim, Kyuhong Shim, Simyung Chang
Comments: Interspeech 2023
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[176] arXiv:2309.00723 (cross-list from cs.CL) [pdf, other]
Title: Contextual Biasing of Named-Entities with Large Language Models
Chuanneng Sun, Zeeshan Ahmed, Yingyi Ma, Zhe Liu, Lucas Kabela, Yutong Pang, Ozlem Kalinli
Comments: 5 pages, 4 figures. Conference: ICASSP 2024
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[177] arXiv:2309.00916 (cross-list from cs.CL) [pdf, html, other]
Title: BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
Chen Wang, Minpeng Liao, Zhongqiang Huang, Jinliang Lu, Junhong Wu, Yuchen Liu, Chengqing Zong, Jiajun Zhang
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[178] arXiv:2309.01076 (cross-list from cs.LG) [pdf, other]
Title: Federated Few-shot Learning for Cough Classification with Edge Devices
Ngan Dao Hoang, Dat Tran-Anh, Manh Luong, Cong Tran, Cuong Pham
Comments: 21 pages, 5 figures
Subjects: Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[179] arXiv:2309.01108 (cross-list from eess.AS) [pdf, html, other]
Title: Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
Sarthak Kumar Maharana, Krishna Kamal Adidam, Shoumik Nandi, Ajitesh Srivastava
Comments: Accepted to IEEE ICASSP Workshops 2024
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[180] arXiv:2309.01142 (cross-list from eess.AS) [pdf, other]
Title: MSM-VC: High-fidelity Source Style Transfer for Non-Parallel Voice Conversion by Multi-scale Style Modeling
Zhichao Wang, Xinsheng Wang, Qicong Xie, Tao Li, Lei Xie, Qiao Tian, Yuping Wang
Comments: This work was submitted on April 10, 2022 and accepted on August 29, 2023
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[181] arXiv:2309.01164 (cross-list from eess.AS) [pdf, other]
Title: Noise robust speech emotion recognition with signal-to-noise ratio adapting speech enhancement
Yu-Wen Chen, Julia Hirschberg, Yu Tsao
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[182] arXiv:2309.01202 (cross-list from cs.GR) [pdf, other]
Title: MAGMA: Music Aligned Generative Motion Autodecoder
Sohan Anisetty, Amit Raj, James Hays
Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[183] arXiv:2309.01513 (cross-list from eess.AS) [pdf, html, other]
Title: RGI-Net: 3D Room Geometry Inference from Room Impulse Responses With Hidden First-Order Reflections
Inmo Yeon, Jung-Woo Choi
Comments: 5 pages, 3 figures, 3 tables
Journal-ref: 2024 18th International Workshop on Acoustic Signal Enhancement (IWAENC), Aalborg, Denmark, 2024, pp. 439-443
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Sound (cs.SD)
[184] arXiv:2309.01535 (cross-list from eess.AS) [pdf, other]
Title: Single-Channel Speech Enhancement with Deep Complex U-Networks and Probabilistic Latent Space Models
Eike J. Nustede, Jörn Anemüller
Journal-ref: ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 2023, pp. 1-5
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[185] arXiv:2309.01576 (cross-list from cs.CL) [pdf, other]
Title: A Comparative Analysis of Pretrained Language Models for Text-to-Speech
Marcel Granero-Moya, Penny Karanasou, Sri Karlapati, Bastian Schnell, Nicole Peinelt, Alexis Moinet, Thomas Drugman
Comments: Accepted for presentation at the 12th ISCA Speech Synthesis Workshop (SSW) in Grenoble, France, from 26th to 28th August 2023
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[186] arXiv:2309.01947 (cross-list from cs.CL) [pdf, other]
Title: TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-device ASR Models
Yuan Shangguan, Haichuan Yang, Danni Li, Chunyang Wu, Yassir Fathullah, Dilin Wang, Ayushi Dalmia, Raghuraman Krishnamoorthi, Ozlem Kalinli, Junteng Jia, Jay Mahadeokar, Xin Lei, Mike Seltzer, Vikas Chandra
Comments: Meta AI; Submitted to ICASSP 2024
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[187] arXiv:2309.01950 (cross-list from cs.CV) [pdf, other]
Title: RADIO: Reference-Agnostic Dubbing Video Synthesis
Dongyeun Lee, Chaewon Kim, Sangjoon Yu, Jaejun Yoo, Gyeong-Moon Park
Comments: Accepted by WACV 2024
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[188] arXiv:2309.02145 (cross-list from cs.CL) [pdf, other]
Title: Bring the Noise: Introducing Noise Robustness to Pretrained Automatic Speech Recognition
Patrick Eickhoff, Matthias Möller, Theresa Pekarek Rosin, Johannes Twiefel, Stefan Wermter
Comments: Submitted and accepted for ICANN 2023 (32nd International Conference on Artificial Neural Networks)
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[189] arXiv:2309.02265 (cross-list from eess.AS) [pdf, other]
Title: PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
Alain Riou, Stefan Lattner, Gaëtan Hadjeres, Geoffroy Peeters
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[190] arXiv:2309.02285 (cross-list from eess.AS) [pdf, other]
Title: PromptTTS 2: Describing and Generating Voices with Text Prompt
Yichong Leng, Zhifang Guo, Kai Shen, Xu Tan, Zeqian Ju, Yanqing Liu, Yufei Liu, Dongchao Yang, Leying Zhang, Kaitao Song, Lei He, Xiang-Yang Li, Sheng Zhao, Tao Qin, Jiang Bian
Comments: Demo page: this https URL
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
[191] arXiv:2309.02405 (cross-list from cs.CV) [pdf, other]
Title: Generating Realistic Images from In-the-wild Sounds
Taegyeong Lee, Jeonghun Kang, Hyeonyu Kim, Taehwan Kim
Comments: Accepted to ICCV 2023
Subjects: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[192] arXiv:2309.02418 (cross-list from eess.AS) [pdf, other]
Title: Personalized Adaptation with Pre-trained Speech Encoders for Continuous Emotion Recognition
Minh Tran, Yufeng Yin, Mohammad Soleymani
Comments: Accepted by INTERSPEECH 2023
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Signal Processing (eess.SP)
[193] arXiv:2309.02432 (cross-list from eess.AS) [pdf, other]
Title: Employing Real Training Data for Deep Noise Suppression
Ziyi Xu, Marvin Sach, Jan Pirklbauer, Tim Fingscheidt
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[194] arXiv:2309.02466 (cross-list from eess.AS) [pdf, other]
Title: Minimal Effective Theory for Phonotactic Memory: Capturing Local Correlations due to Errors in Speech
Paul Myles Eugenio
Comments: 16 pages; 7 figs
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[195] arXiv:2309.02539 (cross-list from eess.AS) [pdf, other]
Title: A Generalized Bandsplit Neural Network for Cinematic Audio Source Separation
Karn N. Watcharasupat, Chih-Wei Wu, Yiwei Ding, Iroro Orife, Aaron J. Hipple, Phillip A. Williams, Scott Kramer, Alexander Lerch, William Wolcott
Comments: Accepted to the IEEE Open Journal of Signal Processing (ICASSP 2024 Track)
Journal-ref: IEEE Open Journal of Signal Processing, vol. 5, pp. 73-81, 2024
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Signal Processing (eess.SP)
[196] arXiv:2309.02567 (cross-list from eess.AS) [pdf, other]
Title: Symbolic Music Representations for Classification Tasks: A Systematic Evaluation
Huan Zhang, Emmanouil Karystinaios, Simon Dixon, Gerhard Widmer, Carlos Eduardo Cancino-Chacón
Comments: To be published in the Proceedings of the 24th International Society for Music Information Retrieval Conference (ISMIR 2023), Milan, Italy
Journal-ref: Proceedings of the 24th International Society for Music Information Retrieval Conference (ISMIR 2023), Milan, Italy
Subjects: Audio and Speech Processing (eess.AS); Multimedia (cs.MM); Sound (cs.SD)
[197] arXiv:2309.02592 (cross-list from eess.AS) [pdf, html, other]
Title: BWSNet: Automatic Perceptual Assessment of Audio Signals
Clément Le Moine Veillon, Victor Rosi, Pablo Arias Sarah, Léane Salais, Nicolas Obin
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[198] arXiv:2309.02730 (cross-list from eess.AS) [pdf, html, other]
Title: Stylebook: Content-Dependent Speaking Style Modeling for Any-to-Any Voice Conversion using Only Speech Data
Hyungseob Lim, Kyungguen Byun, Sunkuk Moon, Erik Visser
Comments: 5 pages, 2 figures, 2 tables
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Sound (cs.SD)
[199] arXiv:2309.02743 (cross-list from eess.AS) [pdf, other]
Title: MuLanTTS: The Microsoft Speech Synthesis System for Blizzard Challenge 2023
Zhihang Xu, Shaofei Zhang, Xi Wang, Jiajun Zhang, Wenning Wei, Lei He, Sheng Zhao
Comments: 6 pages
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[200] arXiv:2309.02780 (cross-list from cs.CL) [pdf, other]
Title: GRASS: Unified Generation Model for Speech-to-Semantic Tasks
Aobo Xia, Shuyu Lei, Yushu Yang, Xiang Guo, Hua Chai
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Total of 451 entries : 1-50 51-100 101-150 151-200 201-250 251-300 301-350 ... 451-451
Showing up to 50 entries per page: fewer | more | all
  • About
  • Help
  • contact arXivClick here to contact arXiv Contact
  • subscribe to arXiv mailingsClick here to subscribe Subscribe
  • Copyright
  • Privacy Policy
  • Web Accessibility Assistance
  • arXiv Operational Status
    Get status notifications via email or slack