Training Recurrent Neural Networks by Diffusion

Mobahi, Hossein

Computer Science > Machine Learning

arXiv:1601.04114 (cs)

[Submitted on 16 Jan 2016 (v1), last revised 4 Feb 2016 (this version, v2)]

Title:Training Recurrent Neural Networks by Diffusion

Authors:Hossein Mobahi

View PDF

Abstract:This work presents a new algorithm for training recurrent neural networks (although ideas are applicable to feedforward networks as well). The algorithm is derived from a theory in nonconvex optimization related to the diffusion equation. The contributions made in this work are two fold. First, we show how some seemingly disconnected mechanisms used in deep learning such as smart initialization, annealed learning rate, layerwise pretraining, and noise injection (as done in dropout and SGD) arise naturally and automatically from this framework, without manually crafting them into the algorithms. Second, we present some preliminary results on comparing the proposed method against SGD. It turns out that the new algorithm can achieve similar level of generalization accuracy of SGD in much fewer number of epochs.

Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:1601.04114 [cs.LG]
	(or arXiv:1601.04114v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1601.04114

Submission history

From: Hossein Mobahi [view email]
[v1] Sat, 16 Jan 2016 02:24:17 UTC (1,131 KB)
[v2] Thu, 4 Feb 2016 23:22:52 UTC (1,132 KB)

Full-text links:

Access Paper:

view license

Current browse context:

< prev | next >

new | recent | 2016-01

Change to browse by:

cs.LG

References & Citations

DBLP - CS Bibliography

listing | bibtex

Hossein Mobahi

export BibTeX citation

Computer Science > Machine Learning

Title:Training Recurrent Neural Networks by Diffusion

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Training Recurrent Neural Networks by Diffusion

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators