Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment
Authors:
Huma Ameer,
Seemab Latif,
Mehwish Fatima
Abstract:
The automated classification of stuttered speech has significant implications for timely assessments providing assistance to speech language pathologists. Despite notable advancements in the field, the cases in which multiple disfluencies occur in speech require attention. We have taken a progressive approach to fill this gap by classifying multi-stuttered speech more efficiently. The problem has…
▽ More
The automated classification of stuttered speech has significant implications for timely assessments providing assistance to speech language pathologists. Despite notable advancements in the field, the cases in which multiple disfluencies occur in speech require attention. We have taken a progressive approach to fill this gap by classifying multi-stuttered speech more efficiently. The problem has been addressed by firstly curating a dataset of multi-stuttered disfluencies from open source dataset SEP-28k audio clips. Secondly, employing Whisper, a state-of-the-art speech recognition model has been leveraged by using its encoder and taking the problem as multi label classification. Thirdly, using a 6 encoder layer Whisper and experimenting with various layer freezing strategies, a computationally efficient configuration of the model was identified. The proposed configuration achieved micro, macro, and weighted F1-scores of 0.88, 0.85, and 0.87, correspondingly on an external test dataset i.e. Fluency-Bank. In addition, through layer freezing strategies, we were able to achieve the aforementioned results by fine-tuning a single encoder layer, consequently, reducing the model's trainable parameters from 20.27 million to 3.29 million. This research study unveils the contribution of the last encoder layer in the identification of disfluencies in stuttered speech. Consequently, it has led to a computationally efficient approach, 83.7% less parameters to train, making the proposed approach more adaptable for various dialects and languages.
△ Less
Submitted 26 February, 2025; v1 submitted 9 June, 2024;
originally announced June 2024.
Dual-Mode Time Domain Multiplexed Chirp Spread Spectrum
Authors:
Ali Waqar Azim,
Ahmad Bazzi,
Mahrukh Fatima,
Raed Shubair,
Marwa Chafii
Abstract:
We propose a dual-mode (DM) time domain multiplexed (TDM) chirp spread spectrum (CSS) modulation for spectral and energy-efficient low-power wide-area networks (LPWANs). DM-CSS modulation that uses both the even and odd cyclic time shifts has been proposed for LPWANs to achieve noteworthy performance improvement over classical counterparts. However, its spectral efficiency (SE) is half of the in-p…
▽ More
We propose a dual-mode (DM) time domain multiplexed (TDM) chirp spread spectrum (CSS) modulation for spectral and energy-efficient low-power wide-area networks (LPWANs). DM-CSS modulation that uses both the even and odd cyclic time shifts has been proposed for LPWANs to achieve noteworthy performance improvement over classical counterparts. However, its spectral efficiency (SE) is half of the in-phase and quadrature (IQ)-TDM-CSS scheme that employs IQ components with both up and down chirps, resulting in a SE that is four times relative to Long Range (LoRa) modulation. Nevertheless, the IQ-TDM-CSS scheme only allows coherent detection. Furthermore, it is also sensitive to carrier frequency and phase offsets, making it less practical for low-cost battery-powered LPWANs for Internet-of-Things (IoT) applications. DM-CSS uses either an up-chirp or a down-chirp. DM-TDM-CSS consists of two chirped symbols that are multiplexed in the time domain. One of these symbols consisting of even and odd frequency shifts (FSs) is chirped using an up-chirp. The second chirped symbol also consists of even and odd FSs, but they are chirped using a down-chirp. It shall be demonstrated that DM-TDM-CSS attains a maximum achievable SE close to IQ-TDM-CSS while also allowing both coherent and non-coherent detection. Additionally, unlike IQ-TDM-CSS, DM-TDM-CSS is robust against carrier frequency and phase offsets.
△ Less
Submitted 8 October, 2022;
originally announced October 2022.