DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input

Rekimoto, Jun

doi:10.1145/3526113.3545685

Computer Science > Human-Computer Interaction

arXiv:2208.10499 (cs)

[Submitted on 22 Aug 2022]

Title:DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input

Authors:Jun Rekimoto

View PDF

Abstract:Interactions based on automatic speech recognition (ASR) have become widely used, with speech input being increasingly utilized to create documents. However, as there is no easy way to distinguish between commands being issued and text required to be input in speech, misrecognitions are difficult to identify and correct, meaning that documents need to be manually edited and corrected. The input of symbols and commands is also challenging because these may be misrecognized as text letters. To address these problems, this study proposes a speech interaction method called DualVoice, by which commands can be input in a whispered voice and letters in a normal voice. The proposed method does not require any specialized hardware other than a regular microphone, enabling a complete hands-free interaction. The method can be used in a wide range of situations where speech recognition is already available, ranging from text input to mobile/wearable computing. Two neural networks were designed in this study, one for discriminating normal speech from whispered speech, and the second for recognizing whisper speech. A prototype of a text input system was then developed to show how normal and whispered voice can be used in speech text input. Other potential applications using DualVoice are also discussed.

Comments:	to appear as ACM UIST 2022 paper
Subjects:	Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
ACM classes:	H.5.2; H.1.2
Cite as:	arXiv:2208.10499 [cs.HC]
	(or arXiv:2208.10499v1 [cs.HC] for this version)
	https://doi.org/10.48550/arXiv.2208.10499
Related DOI:	https://doi.org/10.1145/3526113.3545685

Submission history

From: Jun Rekimoto [view email]
[v1] Mon, 22 Aug 2022 13:01:28 UTC (5,461 KB)

Computer Science > Human-Computer Interaction

Title:DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Human-Computer Interaction

Title:DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators