Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-Agreement

Dou, Zi-Yi; Tu, Zhaopeng; Wang, Xing; Wang, Longyue; Shi, Shuming; Zhang, Tong

Computer Science > Computation and Language

arXiv:1902.05770 (cs)

[Submitted on 15 Feb 2019]

Title:Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-Agreement

Authors:Zi-Yi Dou, Zhaopeng Tu, Xing Wang, Longyue Wang, Shuming Shi, Tong Zhang

View PDF

Abstract:With the promising progress of deep neural networks, layer aggregation has been used to fuse information across layers in various fields, such as computer vision and machine translation. However, most of the previous methods combine layers in a static fashion in that their aggregation strategy is independent of specific hidden states. Inspired by recent progress on capsule networks, in this paper we propose to use routing-by-agreement strategies to aggregate layers dynamically. Specifically, the algorithm learns the probability of a part (individual layer representations) assigned to a whole (aggregated representations) in an iterative way and combines parts accordingly. We implement our algorithm on top of the state-of-the-art neural machine translation model TRANSFORMER and conduct experiments on the widely-used WMT14 English-German and WMT17 Chinese-English translation datasets. Experimental results across language pairs show that the proposed approach consistently outperforms the strong baseline model and a representative static aggregation model.

Comments:	AAAI 2019
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:1902.05770 [cs.CL]
	(or arXiv:1902.05770v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1902.05770

Submission history

From: Zhaopeng Tu [view email]
[v1] Fri, 15 Feb 2019 11:14:35 UTC (5,930 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2019-02

Change to browse by:

cs
cs.AI

References & Citations

DBLP - CS Bibliography

listing | bibtex

Zi-Yi Dou
Zhaopeng Tu
Xing Wang
Longyue Wang
Shuming Shi

…

export BibTeX citation

Computer Science > Computation and Language

Title:Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-Agreement

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-Agreement

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators