Sound

Authors and titles for April 2022

Total of 291 entries : 51-150 101-200 201-291

Showing up to 100 entries per page: fewer | more | all

[51] arXiv:2204.03255 [pdf, other]: Title: Arabic Text-To-Speech (TTS) Data Preparation

Hala Al Masri, Muhy Eddin Za'ter

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[52] arXiv:2204.03307 [pdf, other]: Title: Genre-conditioned Acoustic Models for Automatic Lyrics Transcription of Polyphonic Music

Xiaoxue Gao, Chitralekha Gupta, Haizhou Li

Comments: 5 pages, 1 figure, accepted by IEEE ICASSP 2022

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[53] arXiv:2204.03398 [pdf, other]: Title: Linguistic-Acoustic Similarity Based Accent Shift for Accent Recognition

Qijie Shao, Jinghao Yan, Jian Kang, Pengcheng Guo, Xian Shi, Pengfei Hu, Lei Xie

Comments: Accepted by Interspeech 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[54] arXiv:2204.03421 [pdf, other]: Title: Self-supervised learning for robust voice cloning

Konstantinos Klapsas, Nikolaos Ellinas, Karolos Nikitaras, Georgios Vamvoukakis, Panos Kakoulidis, Konstantinos Markopoulos, Spyros Raptis, June Sig Sung, Gunu Jho, Aimilios Chalamandaris, Pirros Tsiakoulis

Comments: Accepted to INTERSPEECH 2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[55] arXiv:2204.03594 [pdf, other]: Title: Heterogeneous Target Speech Separation

Efthymios Tzinis, Gordon Wichern, Aswin Subramanian, Paris Smaragdis, Jonathan Le Roux

Comments: Submitted to Interspeech 2022

Journal-ref: Interspeech 2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[56] arXiv:2204.03740 [pdf, other]: Title: Successes and critical failures of neural networks in capturing human-like speech recognition

Federico Adolfi, Jeffrey S. Bowers, David Poeppel

Journal-ref: Neural Networks, 162, 199-211 (2023)

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS); Neurons and Cognition (q-bio.NC)
[57] arXiv:2204.03847 [pdf, other]: Title: Enhanced exemplar autoencoder with cycle consistency loss in any-to-one voice conversion

Weida Liang, Lantian Li, Wenqiang Du, Dong Wang

Comments: submitted to INTERSPEECH 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[58] arXiv:2204.03852 [pdf, other]: Title: Reliable Visualization for Deep Speaker Recognition

Pengqi Li, Lantian Li, Askar Hamdulla, Dong Wang

Comments: submitted to INTERSPEECH 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[59] arXiv:2204.03889 [pdf, other]: Title: Adding Connectionist Temporal Summarization into Conformer to Improve Its Decoder Efficiency For Speech Recognition

Nick J.C. Wang, Zongfeng Quan, Shaojun Wang, Jing Xiao

Comments: Submitted to INTERSPEECH 2022 (5 pages, 2 figures)

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[60] arXiv:2204.03967 [pdf, other]: Title: The Sillwood Technologies System for the VoiceMOS Challenge 2022

Jiameng Gao

Comments: Submitted to Interspeech 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[61] arXiv:2204.04166 [pdf, other]: Title: Self-supervised Speaker Diarization

Yehoshua Dissen, Felix Kreuk, Joseph Keshet

Comments: Submitted to Interspeech 2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[62] arXiv:2204.04464 [pdf, other]: Title: Multichannel Speech Separation with Narrow-band Conformer

Changsheng Quan, Xiaofei Li

Comments: accepted by INTERSPEECH 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[63] arXiv:2204.04579 [pdf, other]: Title: Inferring Pitch from Coarse Spectral Features

Danni Ma, Neville Ryant, Mark Liberman

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[64] arXiv:2204.04645 [pdf, other]: Title: Self-Supervised Audio-and-Text Pre-training with Extremely Low-Resource Parallel Data

Yu Kang, Tianqiao Liu, Hang Li, Yang Hao, Wenbiao Ding

Comments: AAAI 2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[65] arXiv:2204.04646 [pdf, other]: Title: Deep Embeddings for Robust User-Based Amateur Vocal Percussion Classification

Alejandro Delgado, Emir Demirel, Vinod Subramanian, Charalampos Saitis, Mark Sandler

Comments: Accepted at Sound and Music Computing (SMC) conference 2022

Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[66] arXiv:2204.04651 [pdf, other]: Title: Deep Conditional Representation Learning for Drum Sample Retrieval by Vocalisation

Alejandro Delgado, Charalampos Saitis, Emmanouil Benetos, Mark Sandler

Comments: Submitted to Interspeech 2022 (under review)

Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[67] arXiv:2204.04756 [pdf, other]: Title: Towards Evaluation of Autonomously Generated Musical Compositions: A Comprehensive Survey

Daniel Kvak

Subjects: Sound (cs.SD); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC); Audio and Speech Processing (eess.AS)
[68] arXiv:2204.04802 [pdf, other]: Title: On the pragmatism of using binary classifiers over data intensive neural network classifiers for detection of COVID-19 from voice

Ankit Shah, Hira Dhamyal, Yang Gao, Daniel Arancibia, Mario Arancibia, Bhiksha Raj, Rita Singh

Comments: Submitted to ICASSP 2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[69] arXiv:2204.04855 [pdf, other]: Title: Fusion of Self-supervised Learned Models for MOS Prediction

Zhengdong Yang, Wangjin Zhou, Chenhui Chu, Sheng Li, Raj Dabre, Raphael Rubino, Yi Zhao

Comments: MOS 2022 shared task system description paper

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[70] arXiv:2204.05070 [pdf, other]: Title: Fine-grained Noise Control for Multispeaker Speech Synthesis

Karolos Nikitaras, Georgios Vamvoukakis, Nikolaos Ellinas, Konstantinos Klapsas, Konstantinos Markopoulos, Spyros Raptis, June Sig Sung, Gunu Jho, Aimilios Chalamandaris, Pirros Tsiakoulis

Comments: Accepted to INTERSPEECH 2022

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[71] arXiv:2204.05082 [pdf, other]: Title: An approach to improving sound-based vehicle speed estimation

Nikola Bulatovic, Slobodan Djukanovic

Comments: Submitted to: 2022 Zooming Innovation in Consumer Technologies Conference (ZINC)

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[72] arXiv:2204.05156 [pdf, other]: Title: How to Listen? Rethinking Visual Sound Localization

Ho-Hsiang Wu, Magdalena Fuentes, Prem Seetharaman, Juan Pablo Bello

Comments: Submitted to INTERSPEECH 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[73] arXiv:2204.05222 [pdf, other]: Title: INTERSPEECH 2022 Audio Deep Packet Loss Concealment Challenge

Lorenz Diener, Sten Sootla, Solomiya Branets, Ando Saabas, Robert Aichner, Ross Cutler

Comments: 4 pages + 1 page references, 1 figure, 2 tables. Submitted to INTERSPEECH 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[74] arXiv:2204.05445 [pdf, other]: Title: Small Footprint Multi-channel ConvMixer for Keyword Spotting with Centroid Based Awareness

Dianwen Ng, Jin Hui Pang, Yang Xiao, Biao Tian, Qiang Fu, Eng Siong Chng

Comments: submitted to INTERSPEECH 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[75] arXiv:2204.05571 [pdf, other]: Title: Speech Emotion Recognition with Global-Aware Fusion on Multi-scale Feature Representation

Wenjing Zhu, Xiang Li

Comments: 6 pages, 3 figures, ICASSP 2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[76] arXiv:2204.05649 [pdf, other]: Title: ADFF: Attention Based Deep Feature Fusion Approach for Music Emotion Recognition

Zi Huang, Shulei Ji, Zhilan Hu, Chuangjian Cai, Jing Luo, Xinyu Yang

Comments: It has been received by Interspeech2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[77] arXiv:2204.06402 [pdf, other]: Title: Sound Event Triage: Detecting Sound Events Considering Priority of Classes

Noriyuki Tonami, Keisuke Imoto

Comments: Accepted to EURASIP Journal on Audio, Speech, and Music Processing

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[78] arXiv:2204.06439 [pdf, other]: Title: Receptive Field Analysis of Temporal Convolutional Networks for Monaural Speech Dereverberation

William Ravenscroft, Stefan Goetze, Thomas Hain

Comments: Accepted at EUSIPCO 2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[79] arXiv:2204.06450 [pdf, other]: Title: The effect of speech pathology on automatic speaker verification -- a large-scale study

Soroosh Tayebi Arasteh, Tobias Weise, Maria Schuster, Elmar Noeth, Andreas Maier, Seung Hee Yang

Comments: Published in Scientific Reports

Journal-ref: Sci Rep 13, 20476 (2023)

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[80] arXiv:2204.06616 [pdf, other]: Title: Predicting score distribution to improve non-intrusive speech quality estimation

Abu Zaher Md Faridee, Hannes Gamper

Comments: Submitted to Interspeech 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[81] arXiv:2204.07018 [pdf, other]: Title: From Environmental Sound Representation to Robustness of 2D CNN Models Against Adversarial Attacks

Mohammad Esmaeilpour, Patrick Cardinal, Alessandro Lameiras Koerich

Comments: 32 pages, Preprint Submitted to Journal of Applied Acoustics. arXiv admin note: substantial text overlap with arXiv:2007.13703

Subjects: Sound (cs.SD); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[82] arXiv:2204.07064 [pdf, other]: Title: Streamable Neural Audio Synthesis With Non-Causal Convolutions

Antoine Caillon, Philippe Esling

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
[83] arXiv:2204.07075 [pdf, other]: Title: Learning and controlling the source-filter representation of speech with a variational autoencoder

Samir Sadok, Simon Leglaive, Laurent Girin, Xavier Alameda-Pineda, Renaud Séguier

Comments: 23 pages, 7 figures, companion website: this https URL

Journal-ref: Speech Communication, vol. 148, 2023

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[84] arXiv:2204.07420 [pdf, other]: Title: Deep CardioSound-An Ensembled Deep Learning Model for Heart Sound MultiLabelling

Li Guo, Steven Davenport, Yonghong Peng

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[85] arXiv:2204.07566 [pdf, other]: Title: Improving Frame-Online Neural Speech Enhancement with Overlapped-Frame Prediction

Zhong-Qiu Wang, Shinji Watanabe

Comments: in IEEE Signal Processing Letters

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[86] arXiv:2204.07763 [pdf, other]: Title: UFRC: A Unified Framework for Reliable COVID-19 Detection on Crowdsourced Cough Audio

Jiangeng Chang, Yucheng Ruan, Cui Shaoze, John Soong Tshon Yit, Mengling Feng

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[87] arXiv:2204.08026 [pdf, other]: Title: Advances in Thunder Sound Synthesis

Eva Fineberg, Jack Walters, Joshua Reiss

Comments: 9 pages, 6 figures, conference paper accepted to the AES Europe Spring 2022 Audio Engineering 152nd Convention

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
[88] arXiv:2204.08164 [pdf, other]: Title: Robust End-to-end Speaker Diarization with Generic Neural Clustering

Chenyu Yang, Yu Wang

Comments: submitted to INTERSPEECH 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[89] arXiv:2204.08269 [pdf, other]: Title: Differentiable Time-Frequency Scattering on GPU

John Muradeli, Cyrus Vahidi, Changhong Wang, Han Han, Vincent Lostanlen, Mathieu Lagrange, George Fazekas

Comments: 8 pages, 6 figures. Submitted to the International Conference on Digital Audio Effects (DAFX) 2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[90] arXiv:2204.08345 [pdf, other]: Title: Extracting Targeted Training Data from ASR Models, and How to Mitigate It

Ehsan Amid, Om Thakkar, Arun Narayanan, Rajiv Mathews, Françoise Beaufays

Comments: Accepted to appear at Interspeech'22

Subjects: Sound (cs.SD); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[91] arXiv:2204.08409 [pdf, other]: Title: Caption Feature Space Regularization for Audio Captioning

Yiming Zhang, Hong Yu, Ruoyi Du, Zhanyu Ma, Yuan Dong

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[92] arXiv:2204.08474 [pdf, other]: Title: AB/BA analysis: A framework for estimating keyword spotting recall improvement while maintaining audio privacy

Raphael Petegrosso, Vasistakrishna Baderdinni, Thibaud Senechal, Benjamin L. Bullough

Comments: Accepted to NAACL 2022 Industry Track

Subjects: Sound (cs.SD); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[93] arXiv:2204.08567 [pdf, other]: Title: Automated Audio Captioning using Audio Event Clues

Ayşegül Özkaya Eren, Mustafa Sert

Comments: submitted to IEEE/ACM Transactions on Audio Speech and Language Processing

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[94] arXiv:2204.08625 [pdf, other]: Title: Self Supervised Adversarial Domain Adaptation for Cross-Corpus and Cross-Language Speech Emotion Recognition

Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, Björn Schuller

Comments: Accepted in IEEE Transactions on Affective Computing

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[95] arXiv:2204.08686 [pdf, other]: Title: Audio-Visual Wake Word Spotting System For MISP Challenge 2021

Yanguang Xu, Jianwei Sun, Yang Han, Shuaijiang Zhao, Chaoyang Mei, Tingwei Guo, Shuran Zhou, Chuandong Xie, Wei Zou, Xiangang Li, Shuran Zhou, Chuandong Xie, Wei Zou, Xiangang Li

Comments: Accepted to ICASSP 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[96] arXiv:2204.08822 [pdf, other]: Title: A Convolutional-Attentional Neural Framework for Structure-Aware Performance-Score Synchronization

Ruchit Agrawal, Daniel Wolff, Simon Dixon

Comments: Published in IEEE Signal Processing Letters, Volume 29, December 2021

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[97] arXiv:2204.08977 [pdf, other]: Title: Disappeared Command: Spoofing Attack On Automatic Speech Recognition Systems with Sound Masking

Jinghui Xu, Jifeng Zhu, Yong Yang

Comments: 13 pages, 4 figures. arXiv admin note: text overlap with arXiv:1903.10346 by other authors

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[98] arXiv:2204.09224 [pdf, other]: Title: ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers

Kaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni, Cheng-I Lai, David Cox, Mark Hasegawa-Johnson, Shiyu Chang

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[99] arXiv:2204.09381 [pdf, other]: Title: Exploration strategies for articulatory synthesis of complex syllable onsets

Daniel R. van Niekerk, Anqi Xu, Branislav Gerazov, Paul K. Krug, Peter Birkholz, Yi Xu

Comments: Accepted at Interspeech 2022

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[100] arXiv:2204.09634 [pdf, other]: Title: Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering

Samuel Lipping, Parthasaarathy Sudarsanam, Konstantinos Drossos, Tuomas Virtanen

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[101] arXiv:2204.09883 [pdf, other]: Title: Layer-wise Fast Adaptation for End-to-End Multi-Accent Speech Recognition

Xun Gong, Yizhou Lu, Zhikai Zhou, Yanmin Qian

Comments: Accepted by Interspeech2021

Journal-ref: Proc. Interspeech 2021

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[102] arXiv:2204.09911 [pdf, other]: Title: STFT-Domain Neural Speech Enhancement with Very Low Algorithmic Latency

Zhong-Qiu Wang, Gordon Wichern, Shinji Watanabe, Jonathan Le Roux

Comments: in IEEE/ACM Transactions on Audio, Speech, and Language Processing

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[103] arXiv:2204.09917 [pdf, other]: Title: SinTra: Learning an inspiration model from a single multi-track music segment

Qingwei Song, Qiwei Sun, Dongsheng Guo, Haiyong Zheng

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[104] arXiv:2204.09976 [pdf, other]: Title: Baseline Systems for the First Spoofing-Aware Speaker Verification Challenge: Score and Embedding Fusion

Hye-jin Shim, Hemlata Tak, Xuechen Liu, Hee-Soo Heo, Jee-weon Jung, Joon Son Chung, Soo-Whan Chung, Ha-Jin Yu, Bong-Jin Lee, Massimiliano Todisco, Héctor Delgado, Kong Aik Lee, Md Sahidullah, Tomi Kinnunen, Nicholas Evans

Comments: 8 pages, accepted by Odyssey 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[105] arXiv:2204.10125 [pdf, other]: Title: Physical Modeling using Recurrent Neural Networks with Fast Convolutional Layers

Julian D. Parker, Sebastian J. Schlecht, Rudolf Rabenstein, Maximilian Schäfer

Comments: Accepted to DAFx2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Computational Physics (physics.comp-ph)
[106] arXiv:2204.10523 [pdf, other]: Title: Unifying Cosine and PLDA Back-ends for Speaker Verification

Zhiyuan Peng, Xuanji He, Ke Ding, Tan Lee, Guanglu Wan

Comments: submitted to interspeech2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[107] arXiv:2204.10561 [pdf, other]: Title: Speaking-Rate-Controllable HiFi-GAN Using Feature Interpolation

Detai Xin, Shinnosuke Takamichi, Takuma Okamoto, Hisashi Kawai, Hiroshi Saruwatari

Comments: submitted to INTERSPEECH 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[108] arXiv:2204.10581 [pdf, other]: Title: Fused Audio Instance and Representation for Respiratory Disease Detection

Tuan Truong, Matthias Lenga, Antoine Serrurier, Sadegh Mohammadi

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[109] arXiv:2204.10749 [pdf, other]: Title: E2E Segmenter: Joint Segmenting and Decoding for Long-Form ASR

W. Ronny Huang, Shuo-yiin Chang, David Rybach, Rohit Prabhavalkar, Tara N. Sainath, Cyril Allauzen, Cal Peyser, Zhiyun Lu

Comments: Interspeech 2022

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[110] arXiv:2204.11139 [pdf, other]: Title: Musical Stylistic Analysis: A Study of Intervallic Transition Graphs via Persistent Homology

Martín Mijangos, Alessandro Bravetti, Pablo Padilla

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Algebraic Topology (math.AT)
[111] arXiv:2204.11304 [pdf, other]: Title: Dictionary Attacks on Speaker Verification

Mirko Marras, Pawel Korus, Anubhav Jain, Nasir Memon

Comments: Accepted in IEEE Transactions on Information Forensics and Security

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[112] arXiv:2204.11320 [pdf, other]: Title: Emotion-Aware Transformer Encoder for Empathetic Dialogue Generation

Raman Goel, Seba Susan, Sachin Vashisht, Armaan Dhanda

Comments: Accepted in 2021 9th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW)

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[113] arXiv:2204.11382 [pdf, other]: Title: Real-time Speech Emotion Recognition Based on Syllable-Level Feature Extraction

Abdul Rehman, Zhen-Tao Liu, Min Wu, Wei-Hua Cao, Cheng-Shan Jiang

Comments: Significant revisions

Subjects: Sound (cs.SD); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[114] arXiv:2204.11403 [pdf, other]: Title: Back-ends Selection for Deep Speaker Embeddings

Zhuo Li, Runqiu Xiao, Zihan Zhang, Zhenduo Zhao, Wenchao Wang, Pengyuan Zhang

Comments: submitted to interspeech2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[115] arXiv:2204.11437 [pdf, other]: Title: Understanding Audio Features via Trainable Basis Functions

Kwan Yee Heung, Kin Wai Cheuk, Dorien Herremans

Comments: under review in Interspeech 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
[116] arXiv:2204.11479 [pdf, other]: Title: End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network

Avi Gazneli, Gadi Zimerman, Tal Ridnik, Gilad Sharir, Asaf Noy

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[117] arXiv:2204.11792 [pdf, other]: Title: SyntaSpeech: Syntax-Aware Generative Adversarial Text-to-Speech

Zhenhui Ye, Zhou Zhao, Yi Ren, Fei Wu

Comments: Accepted by IJCAI-2022. 12 pages

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[118] arXiv:2204.11806 [pdf, html, other]: Title: Parallel Synthesis for Autoregressive Speech Generation

Po-chun Hsu, Da-rong Liu, Andy T. Liu, Hung-yi Lee

Comments: IEEE/ACM Transactions on Audio, Speech, and Language Processing

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[119] arXiv:2204.11942 [pdf, other]: Title: Meta-AF: Meta-Learning for Adaptive Filters

Jonah Casebeer, Nicholas J. Bryan, Paris Smaragdis

Comments: Accepted to ACM/IEEE TASLP. Source code and audio examples: this https URL

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
[120] arXiv:2204.12112 [pdf, other]: Title: Reformulating Speaker Diarization as Community Detection With Emphasis On Topological Structure

Siqi Zheng, Hongbin Suo

Comments: ICASSP 2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[121] arXiv:2204.12177 [pdf, other]: Title: A Comparative Study on Approaches to Acoustic Scene Classification using CNNs

Ishrat Jahan Ananya, Sarah Suad, Shadab Hafiz Choudhury, Mohammad Ashrafuzzaman Khan

Comments: Presented at 2021 Mexican International Conference on Artificial Intelligence. Published in Advances in Computational Intelligence, MICAI 2021, Lecture Notes in Computer Science. 12 pages, 3 figures, 5 tables

Journal-ref: Advances in Computational Intelligence, MICAI 2021, Lecture Notes in Artificial Intelligence vol. 13067, pp. 81-91 (2021)

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[122] arXiv:2204.12290 [pdf, other]: Title: On Machine Learning-Driven Surrogates for Sound Transmission Loss Simulations

Barbara Cunha (LTDS), Abdel-Malek Zine (ICJ), Mohamed Ichchou (ECL), Christophe Droz (COSYS-SII), Stéphane Foulard

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Medical Physics (physics.med-ph)
[123] arXiv:2204.12486 [pdf, other]: Title: Measurement uncertainty and unicity of single number quantities describing the spatial decay of speech level in open-plan offices

Lucas Lenne (INRS (Vandoeuvre lès Nancy)), Patrick Chevret, Étienne Parizet

Journal-ref: Applied Acoustics, Elsevier, 2021, 182, pp.108269

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[124] arXiv:2204.12622 [pdf, other]: Title: Named Entity Recognition for Audio De-Identification

Guillaume Baril, Patrick Cardinal, Alessandro Lameiras Koerich

Comments: 8 pages

Subjects: Sound (cs.SD); Cryptography and Security (cs.CR); Audio and Speech Processing (eess.AS)
[125] arXiv:2204.12768 [pdf, other]: Title: Masked Spectrogram Prediction For Self-Supervised Audio Pre-Training

Dading Chong, Helin Wang, Peilin Zhou, Qingcheng Zeng

Comments: Submit to INTERSPEECH 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[126] arXiv:2204.13094 [pdf, other]: Title: Unsupervised Word Segmentation using K Nearest Neighbors

Tzeviya Sylvia Fuchs, Yedid Hoshen, Joseph Keshet

Comments: Submitted to interspeech 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[127] arXiv:2204.13206 [pdf, other]: Title: Improving Multimodal Speech Recognition by Data Augmentation and Speech Representations

Dan Oneata, Horia Cucu

Comments: Accepted at the Multimodal Learning and Applications Workshop (MULA) from CVPR 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)
[128] arXiv:2204.13289 [pdf, other]: Title: Music Enhancement via Image Translation and Vocoding

Nikhil Kandpal, Oriol Nieto, Zeyu Jin

Comments: ICASSP 2022

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[129] arXiv:2204.13430 [pdf, other]: Title: Pseudo strong labels for large scale weakly supervised audio tagging

Heinrich Dinkel, Zhiyong Yan, Yongqing Wang, Junbo Zhang, Yujun Wang

Comments: Accepted by ICASSP 2022

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[130] arXiv:2204.13437 [pdf, other]: Title: Regotron: Regularizing the Tacotron2 architecture via monotonic alignment loss

Efthymios Georgiou, Kosmas Kritsis, Georgios Paraskevopoulos, Athanasios Katsamanis, Vassilis Katsouros, Alexandros Potamianos

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[131] arXiv:2204.13601 [pdf, other]: Title: Emotion Recognition In Persian Speech Using Deep Neural Networks

Ali Yazdani, Hossein Simchi, Yasser Shekofteh

Comments: 5 pages, 1 figure, 3 tables

Journal-ref: 11th International Conference on Computer and Knowledge Engineering (ICCKE 2021)

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[132] arXiv:2204.13668 [pdf, other]: Title: Unaligned Supervision For Automatic Music Transcription in The Wild

Ben Maman, Amit H. Bermano

Comments: 16 pages, project page available at this https URL

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[133] arXiv:2204.14057 [pdf, other]: Title: Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast

Boqing Zhu, Kele Xu, Changjian Wang, Zheng Qin, Tao Sun, Huaimin Wang, Yuxing Peng

Comments: 8 pages, 4 figures. Accepted by IJCAI-2022

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[134] arXiv:2204.00065 (cross-list from eess.AS) [pdf, other]: Title: Importance of Different Temporal Modulations of Speech: A Tale of Two Perspectives

Samik Sadhu, Hynek Hermansky

Comments: Submitted to ICASSP 2023

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[135] arXiv:2204.00164 (cross-list from cs.CL) [pdf, other]: Title: Filter-based Discriminative Autoencoders for Children Speech Recognition

Chiang-Lin Tai, Hung-Shin Lee, Yu Tsao, Hsin-Min Wang

Comments: Published in EUSIPCO 2022

Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[136] arXiv:2204.00170 (cross-list from eess.AS) [pdf, other]: Title: Universal Adaptor: Converting Mel-Spectrograms Between Different Configurations for Speech Synthesis

Fan-Lin Wang, Po-chun Hsu, Da-rong Liu, Hung-yi Lee

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[137] arXiv:2204.00174 (cross-list from cs.CL) [pdf, other]: Title: InterAug: Augmenting Noisy Intermediate Predictions for CTC-based ASR

Yu Nakagome, Tatsuya Komatsu, Yusuke Fujita, Shuta Ichimura, Yusuke Kida

Comments: This paper was submitted to INTERSPEECH2022

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[138] arXiv:2204.00175 (cross-list from cs.CL) [pdf, other]: Title: Alternate Intermediate Conditioning with Syllable-level and Character-level Targets for Japanese ASR

Yusuke Fujita, Tatsuya Komatsu, Yusuke Kida

Comments: SLT 2022

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[139] arXiv:2204.00176 (cross-list from cs.CL) [pdf, other]: Title: Better Intermediates Improve CTC Inference

Tatsuya Komatsu, Yusuke Fujita, Jaesong Lee, Lukas Lee, Shinji Watanabe, Yusuke Kida

Comments: 5 pages, submitted INTERSPEECH2022

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[140] arXiv:2204.00212 (cross-list from cs.CL) [pdf, other]: Title: Effect and Analysis of Large-scale Language Model Rescoring on Competitive ASR Systems

Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Nobuyasu Itoh, George Saon

Comments: Accepted to Interspeech 2022

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[141] arXiv:2204.00218 (cross-list from eess.AS) [pdf, other]: Title: End-to-End Multi-speaker ASR with Independent Vector Analysis

Robin Scheibler, Wangyou Zhang, Xuankai Chang, Shinji Watanabe, Yanmin Qian

Comments: Submitted to INTERSPEECH2022. 5 pages, 2 figures, 3 tables

Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[142] arXiv:2204.00291 (cross-list from cs.CL) [pdf, other]: Title: Text-To-Speech Data Augmentation for Low Resource Speech Recognition

Rodolfo Zevallos

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[143] arXiv:2204.00348 (cross-list from cs.CL) [pdf, other]: Title: WavFT: Acoustic model finetuning with labelled and unlabelled data

Utkarsh Chauhan, Vikas Joshi, Rupesh R. Mehta

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[144] arXiv:2204.00436 (cross-list from eess.AS) [pdf, other]: Title: AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios

Yihan Wu, Xu Tan, Bohan Li, Lei He, Sheng Zhao, Ruihua Song, Tao Qin, Tie-Yan Liu

Comments: 5 pages, 2 tables, 2 figure. Submitted to Interspeech 2022

Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)
[145] arXiv:2204.00555 (cross-list from eess.AS) [pdf, other]: Title: 1-D CNN based Acoustic Scene Classification via Reducing Layer-wise Dimensionality

Arshdeep Singh

Comments: No comments

Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[146] arXiv:2204.00558 (cross-list from cs.CL) [pdf, other]: Title: Multi-task RNN-T with Semantic Decoder for Streamable Spoken Language Understanding

Xuandi Fu, Feng-Ju Chang, Martin Radfar, Kai Wei, Jing Liu, Grant P. Strimel, Kanthashree Mysore Sathyendra

Comments: Accepted at ICASSP 2022

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[147] arXiv:2204.00604 (cross-list from cs.CV) [pdf, other]: Title: Quantized GAN for Complex Music Generation from Dance Videos

Ye Zhu, Kyle Olszewski, Yu Wu, Panos Achlioptas, Menglei Chai, Yan Yan, Sergey Tulyakov

Comments: Dataset and code at this https URL

Subjects: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[148] arXiv:2204.00618 (cross-list from eess.AS) [pdf, other]: Title: ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion

Edresson Casanova, Christopher Shulby, Alexander Korolev, Arnaldo Candido Junior, Anderson da Silva Soares, Sandra Aluísio, Moacir Antonelli Ponti

Comments: This paper was accepted at INTERSPEECH 2023

Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[149] arXiv:2204.00657 (cross-list from eess.AS) [pdf, other]: Title: Multimodal Clustering with Role Induced Constraints for Speaker Diarization

Nikolaos Flemotomos, Shrikanth Narayanan

Comments: To appear at Interspeech 2022

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[150] arXiv:2204.00679 (cross-list from cs.CV) [pdf, other]: Title: Learning Audio-Video Modalities from Image Captions

Arsha Nagrani, Paul Hongsuck Seo, Bryan Seybold, Anja Hauth, Santiago Manen, Chen Sun, Cordelia Schmid

Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)

Total of 291 entries : 51-150 101-200 201-291

Showing up to 100 entries per page: fewer | more | all