Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODE

Zhuang, Juntang; Dvornek, Nicha; Li, Xiaoxiao; Tatikonda, Sekhar; Papademetris, Xenophon; Duncan, James

Statistics > Machine Learning

arXiv:2006.02493 (stat)

[Submitted on 3 Jun 2020]

Title:Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODE

Authors:Juntang Zhuang, Nicha Dvornek, Xiaoxiao Li, Sekhar Tatikonda, Xenophon Papademetris, James Duncan

View PDF

Abstract:Neural ordinary differential equations (NODEs) have recently attracted increasing attention; however, their empirical performance on benchmark tasks (e.g. image classification) are significantly inferior to discrete-layer models. We demonstrate an explanation for their poorer performance is the inaccuracy of existing gradient estimation methods: the adjoint method has numerical errors in reverse-mode integration; the naive method directly back-propagates through ODE solvers, but suffers from a redundantly deep computation graph when searching for the optimal stepsize. We propose the Adaptive Checkpoint Adjoint (ACA) method: in automatic differentiation, ACA applies a trajectory checkpoint strategy which records the forward-mode trajectory as the reverse-mode trajectory to guarantee accuracy; ACA deletes redundant components for shallow computation graphs; and ACA supports adaptive solvers. On image classification tasks, compared with the adjoint and naive method, ACA achieves half the error rate in half the training time; NODE trained with ACA outperforms ResNet in both accuracy and test-retest reliability. On time-series modeling, ACA outperforms competing methods. Finally, in an example of the three-body problem, we show NODE with ACA can incorporate physical knowledge to achieve better accuracy. We provide the PyTorch implementation of ACA: \url{this https URL}.

Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as:	arXiv:2006.02493 [stat.ML]
	(or arXiv:2006.02493v1 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2006.02493
Journal reference:	https://proceedings.icml.cc/static/paper_files/icml/2020/917-Paper.pdf

Submission history

From: Juntang Zhuang [view email]
[v1] Wed, 3 Jun 2020 19:48:00 UTC (2,503 KB)

Statistics > Machine Learning

Title:Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODE

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODE

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators