Learning Online Alignments with Continuous Rewards Policy Gradient

Luo, Yuping; Chiu, Chung-Cheng; Jaitly, Navdeep; Sutskever, Ilya

Computer Science > Machine Learning

arXiv:1608.01281 (cs)

[Submitted on 3 Aug 2016]

Title:Learning Online Alignments with Continuous Rewards Policy Gradient

Authors:Yuping Luo, Chung-Cheng Chiu, Navdeep Jaitly, Ilya Sutskever

View PDF

Abstract:Sequence-to-sequence models with soft attention had significant success in machine translation, speech recognition, and question answering. Though capable and easy to use, they require that the entirety of the input sequence is available at the beginning of inference, an assumption that is not valid for instantaneous translation and speech recognition. To address this problem, we present a new method for solving sequence-to-sequence problems using hard online alignments instead of soft offline alignments. The online alignments model is able to start producing outputs without the need to first process the entire input sequence. A highly accurate online sequence-to-sequence model is useful because it can be used to build an accurate voice-based instantaneous translator. Our model uses hard binary stochastic decisions to select the timesteps at which outputs will be produced. The model is trained to produce these stochastic decisions using a standard policy gradient method. In our experiments, we show that this model achieves encouraging performance on TIMIT and Wall Street Journal (WSJ) speech recognition datasets.

Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as:	arXiv:1608.01281 [cs.LG]
	(or arXiv:1608.01281v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1608.01281

Submission history

From: Navdeep Jaitly [view email]
[v1] Wed, 3 Aug 2016 18:35:12 UTC (451 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2016-08

Change to browse by:

cs
cs.CL

References & Citations

DBLP - CS Bibliography

listing | bibtex

Yuping Luo
Chung-Cheng Chiu
Navdeep Jaitly
Ilya Sutskever

export BibTeX citation

Computer Science > Machine Learning

Title:Learning Online Alignments with Continuous Rewards Policy Gradient

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Learning Online Alignments with Continuous Rewards Policy Gradient

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators