Skip to main content

Showing 1–11 of 11 results for author: Coler, M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2506.02627  [pdf, ps, other

    cs.CL cs.SD eess.AS

    Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning

    Authors: Ömer Tarik Özyilmaz, Matt Coler, Matias Valdenegro-Toro

    Abstract: Although commercial Arabic automatic speech recognition (ASR) systems support Modern Standard Arabic (MSA), they struggle with dialectal speech. We investigate the effect of fine-tuning OpenAI's Whisper on five major Arabic dialects (Gulf, Levantine, Iraqi, Egyptian, Maghrebi) using Mozilla Common Voice for MSA and the MASC dataset for dialectal speech. We evaluate MSA training size effects, benef… ▽ More

    Submitted 3 June, 2025; originally announced June 2025.

    Comments: Accepted at Interspeech 2025

  2. arXiv:2506.00955  [pdf, ps, other

    cs.CL cs.SD eess.AS

    Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection

    Authors: Zhu Li, Yuqing Zhang, Xiyuan Gao, Shekhar Nayak, Matt Coler

    Abstract: Sarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, existing detection systems often rely on multimodal data, limiting their applicability in contexts where only speech is available. To address this, we propose an annotation pipeline that leverages large language models (LLMs) to generate a sarcasm dataset… ▽ More

    Submitted 1 June, 2025; originally announced June 2025.

    Comments: Accepted to Interspeech 2025

  3. arXiv:2502.04883  [pdf, other

    cs.CL cs.LG cs.SD eess.AS

    Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance

    Authors: Reihaneh Amooie, Wietse de Vries, Yun Hao, Jelske Dijkstra, Matt Coler, Martijn Wieling

    Abstract: Automatic Speech Recognition (ASR) performance for low-resource languages is still far behind that of higher-resource languages such as English, due to a lack of sufficient labeled data. State-of-the-art methods deploy self-supervised transfer learning where a model pre-trained on large amounts of data is fine-tuned using little labeled data in a target low-resource language. In this paper, we pre… ▽ More

    Submitted 7 February, 2025; originally announced February 2025.

  4. arXiv:2412.10103  [pdf, other

    cs.CL

    AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation

    Authors: Xiyuan Gao, Shubhi Bansal, Kushaan Gowda, Zhu Li, Shekhar Nayak, Nagendra Kumar, Matt Coler

    Abstract: Detecting sarcasm effectively requires a nuanced understanding of context, including vocal tones and facial expressions. The progression towards multimodal computational methods in sarcasm detection, however, faces challenges due to the scarcity of data. To address this, we present AMuSeD (Attentive deep neural network for MUltimodal Sarcasm dEtection incorporating bi-modal Data augmentation). Thi… ▽ More

    Submitted 13 December, 2024; originally announced December 2024.

    Comments: This is a preprint version of the paper, submitted and under review at the IEEE Transactions on Affective Computing

  5. arXiv:2408.14892  [pdf, other

    cs.CL cs.SD eess.AS

    A Functional Trade-off between Prosodic and Semantic Cues in Conveying Sarcasm

    Authors: Zhu Li, Xiyuan Gao, Yuqing Zhang, Shekhar Nayak, Matt Coler

    Abstract: This study investigates the acoustic features of sarcasm and disentangles the interplay between the propensity of an utterance being used sarcastically and the presence of prosodic cues signaling sarcasm. Using a dataset of sarcastic utterances compiled from television shows, we analyze the prosodic features within utterances and key phrases belonging to three distinct sarcasm categories (embedded… ▽ More

    Submitted 27 August, 2024; originally announced August 2024.

    Comments: accepted at Interspeech 2024

  6. arXiv:2406.06403  [pdf, other

    cs.CL cs.LG cs.SD eess.AS

    Meta Learning Text-to-Speech Synthesis in over 7000 Languages

    Authors: Florian Lux, Sarina Meyer, Lyonel Behringer, Frank Zalkow, Phat Do, Matt Coler, Emanuël A. P. Habets, Ngoc Thang Vu

    Abstract: In this work, we take on the challenging task of building a single text-to-speech synthesis system that is capable of generating speech in over 7000 languages, many of which lack sufficient data for traditional TTS development. By leveraging a novel integration of massively multilingual pretraining and meta learning to approximate language representations, our approach enables zero-shot speech syn… ▽ More

    Submitted 10 June, 2024; originally announced June 2024.

    Comments: accepted at Interspeech 2024

  7. arXiv:2306.12040  [pdf, other

    cs.CL eess.AS

    Strategies in Transfer Learning for Low-Resource Speech Synthesis: Phone Mapping, Features Input, and Source Language Selection

    Authors: Phat Do, Matt Coler, Jelske Dijkstra, Esther Klabbers

    Abstract: We compare using a PHOIBLE-based phone mapping method and using phonological features input in transfer learning for TTS in low-resource languages. We use diverse source languages (English, Finnish, Hindi, Japanese, and Russian) and target languages (Bulgarian, Georgian, Kazakh, Swahili, Urdu, and Uzbek) to test the language-independence of the methods and enhance the findings' applicability. We u… ▽ More

    Submitted 21 June, 2023; originally announced June 2023.

    Comments: Accepted at the Speech Synthesis Workshop 2023

  8. arXiv:2306.00535  [pdf, other

    cs.CL eess.AS

    The Effects of Input Type and Pronunciation Dictionary Usage in Transfer Learning for Low-Resource Text-to-Speech

    Authors: Phat Do, Matt Coler, Jelske Dijkstra, Esther Klabbers

    Abstract: We compare phone labels and articulatory features as input for cross-lingual transfer learning in text-to-speech (TTS) for low-resource languages (LRLs). Experiments with FastSpeech 2 and the LRL West Frisian show that using articulatory features outperformed using phone labels in both intelligibility and naturalness. For LRLs without pronunciation dictionaries, we propose two novel approaches: a)… ▽ More

    Submitted 1 June, 2023; originally announced June 2023.

    Comments: Accepted at INTERSPEECH 2023

  9. arXiv:2305.19396  [pdf, other

    eess.AS cs.CL

    Resource-Efficient Fine-Tuning Strategies for Automatic MOS Prediction in Text-to-Speech for Low-Resource Languages

    Authors: Phat Do, Matt Coler, Jelske Dijkstra, Esther Klabbers

    Abstract: We train a MOS prediction model based on wav2vec 2.0 using the open-access data sets BVCC and SOMOS. Our test with neural TTS data in the low-resource language (LRL) West Frisian shows that pre-training on BVCC before fine-tuning on SOMOS leads to the best accuracy for both fine-tuned and zero-shot prediction. Further fine-tuning experiments show that using more than 30 percent of the total data d… ▽ More

    Submitted 30 May, 2023; originally announced May 2023.

    Comments: Accepted at INTERSPEECH 2023

  10. arXiv:2205.03608  [pdf, other

    cs.CL

    UniMorph 4.0: Universal Morphology

    Authors: Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina J. Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Benoît Sagot, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay , et al. (71 additional authors not shown)

    Abstract: The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. This pa… ▽ More

    Submitted 19 June, 2022; v1 submitted 7 May, 2022; originally announced May 2022.

    Comments: LREC 2022; The first two authors made equal contributions

  11. Evolving Plasticity for Autonomous Learning under Changing Environmental Conditions

    Authors: Anil Yaman, Giovanni Iacca, Decebal Constantin Mocanu, Matt Coler, George Fletcher, Mykola Pechenizkiy

    Abstract: A fundamental aspect of learning in biological neural networks is the plasticity property which allows them to modify their configurations during their lifetime. Hebbian learning is a biologically plausible mechanism for modeling the plasticity property in artificial neural networks (ANNs), based on the local interactions of neurons. However, the emergence of a coherent global learning behavior fr… ▽ More

    Submitted 7 December, 2020; v1 submitted 2 April, 2019; originally announced April 2019.

    Comments: Evolutionary Computation Journal

    Journal ref: Evolutionary Computation 1 25, 2020