oViT: An Accurate Second-Order Pruning Framework for Vision Transformers

Kuznedelev, Denis; Kurtic, Eldar; Frantar, Elias; Alistarh, Dan

Computer Science > Computer Vision and Pattern Recognition

arXiv:2210.09223v1 (cs)

[Submitted on 14 Oct 2022 (this version), latest version 31 May 2023 (v2)]

Title:oViT: An Accurate Second-Order Pruning Framework for Vision Transformers

Authors:Denis Kuznedelev, Eldar Kurtic, Elias Frantar, Dan Alistarh

View PDF

Abstract:Models from the Vision Transformer (ViT) family have recently provided breakthrough results across image classification tasks such as ImageNet. Yet, they still face barriers to deployment, notably the fact that their accuracy can be severely impacted by compression techniques such as pruning. In this paper, we take a step towards addressing this issue by introducing Optimal ViT Surgeon (oViT), a new state-of-the-art method for the weight sparsification of Vision Transformers (ViT) models. At the technical level, oViT introduces a new weight pruning algorithm which leverages second-order information, specifically adapted to be both highly-accurate and efficient in the context of ViTs. We complement this accurate one-shot pruner with an in-depth investigation of gradual pruning, augmentation, and recovery schedules for ViTs, which we show to be critical for successful ViT compression. We validate our method via extensive experiments on classical ViT and DeiT models, as well as on newer variants, such as XCiT, EfficientFormer and Swin. Moreover, our results are even relevant to recently-proposed highly-accurate ResNets. Our results show for the first time that ViT-family models can in fact be pruned to high sparsity levels (e.g. $\geq 75\%$) with low impact on accuracy ($\leq 1\%$ relative drop), and that our approach outperforms prior methods by significant margins at high sparsities. In addition, we show that our method is compatible with structured pruning methods and quantization, and that it can lead to significant speedups on a sparsity-aware inference engine.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
MSC classes:	68T07
ACM classes:	I.m
Cite as:	arXiv:2210.09223 [cs.CV]
	(or arXiv:2210.09223v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2210.09223

Submission history

From: Denis Kuznedelev [view email]
[v1] Fri, 14 Oct 2022 12:19:09 UTC (3,368 KB)
[v2] Wed, 31 May 2023 09:59:46 UTC (5,026 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:oViT: An Accurate Second-Order Pruning Framework for Vision Transformers

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:oViT: An Accurate Second-Order Pruning Framework for Vision Transformers

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators