Initial Decoding with Minimally Augmented Language Model for Improved Lattice Rescoring in Low Resource ASR

Murthy, Savitha; Sitaram, Dinkar

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2403.10937 (eess)

[Submitted on 16 Mar 2024]

Title:Initial Decoding with Minimally Augmented Language Model for Improved Lattice Rescoring in Low Resource ASR

Authors:Savitha Murthy, Dinkar Sitaram

View PDF HTML (experimental)

Abstract:This paper addresses the problem of improving speech recognition accuracy with lattice rescoring in low-resource languages where the baseline language model is insufficient for generating inclusive lattices. We minimally augment the baseline language model with word unigram counts that are present in a larger text corpus of the target language but absent in the baseline. The lattices generated after decoding with such an augmented baseline language model are more comprehensive. We obtain 21.8% (Telugu) and 41.8% (Kannada) relative word error reduction with our proposed method. This reduction in word error rate is comparable to 21.5% (Telugu) and 45.9% (Kannada) relative word error reduction obtained by decoding with full Wikipedia text augmented language mode while our approach consumes only 1/8th the memory. We demonstrate that our method is comparable with various text selection-based language model augmentation and also consistent for data sets of different sizes. Our approach is applicable for training speech recognition systems under low resource conditions where speech data and compute resources are insufficient, while there is a large text corpus that is available in the target language. Our research involves addressing the issue of out-of-vocabulary words of the baseline in general and does not focus on resolving the absence of named entities. Our proposed method is simple and yet computationally less expensive.

Comments:	14 pages, 7 figures, Accepted in Sadhana Journal
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2403.10937 [eess.AS]
	(or arXiv:2403.10937v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2403.10937

Submission history

From: Savitha Murthy Ms. [view email]
[v1] Sat, 16 Mar 2024 14:34:31 UTC (1,481 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Initial Decoding with Minimally Augmented Language Model for Improved Lattice Rescoring in Low Resource ASR

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Initial Decoding with Minimally Augmented Language Model for Improved Lattice Rescoring in Low Resource ASR

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators