Nonwords Pronunciation Classification in Language Development Tests for Preschool Children

Baumann, Ilja; Wagner, Dominik; Bayerl, Sebastian; Bocklet, Tobias

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2206.08058 (eess)

[Submitted on 16 Jun 2022 (v1), last revised 17 Jun 2022 (this version, v2)]

Title:Nonwords Pronunciation Classification in Language Development Tests for Preschool Children

Authors:Ilja Baumann, Dominik Wagner, Sebastian Bayerl, Tobias Bocklet

View PDF

Abstract:This work aims to automatically evaluate whether the language development of children is age-appropriate. Validated speech and language tests are used for this purpose to test the auditory memory. In this work, the task is to determine whether spoken nonwords have been uttered correctly. We compare different approaches that are motivated to model specific language structures: Low-level features (FFT), speaker embeddings (ECAPA-TDNN), grapheme-motivated embeddings (wav2vec 2.0), and phonetic embeddings in form of senones (ASR acoustic model). Each of the approaches provides input for VGG-like 5-layer CNN classifiers. We also examine the adaptation per nonword. The evaluation of the proposed systems was performed using recordings from different kindergartens of spoken nonwords. ECAPA-TDNN and low-level FFT features do not explicitly model phonetic information; wav2vec2.0 is trained on grapheme labels, our ASR acoustic model features contain (sub-)phonetic information. We found that the more granular the phonetic modeling is, the higher are the achieved recognition rates. The best system trained on ASR acoustic model features with VTLN achieved an accuracy of 89.4% and an area under the ROC (Receiver Operating Characteristic) curve (AUC) of 0.923. This corresponds to an improvement in accuracy of 20.2% and AUC of 0.309 relative compared to the FFT-baseline.

Comments:	Accepted at Interspeech 2022
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
Cite as:	arXiv:2206.08058 [eess.AS]
	(or arXiv:2206.08058v2 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2206.08058

Submission history

From: Ilja Baumann [view email]
[v1] Thu, 16 Jun 2022 10:19:47 UTC (718 KB)
[v2] Fri, 17 Jun 2022 07:08:27 UTC (718 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Nonwords Pronunciation Classification in Language Development Tests for Preschool Children

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Nonwords Pronunciation Classification in Language Development Tests for Preschool Children

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators