Instant Quantization of Neural Networks using Monte Carlo Methods

Mordido, Gonçalo; Van Keirsbilck, Matthijs; Keller, Alexander

Computer Science > Machine Learning

arXiv:1905.12253 (cs)

[Submitted on 29 May 2019 (v1), last revised 7 Jan 2020 (this version, v2)]

Title:Instant Quantization of Neural Networks using Monte Carlo Methods

Authors:Gonçalo Mordido, Matthijs Van Keirsbilck, Alexander Keller

View PDF

Abstract:Low bit-width integer weights and activations are very important for efficient inference, especially with respect to lower power consumption. We propose Monte Carlo methods to quantize the weights and activations of pre-trained neural networks without any re-training. By performing importance sampling we obtain quantized low bit-width integer values from full-precision weights and activations. The precision, sparsity, and complexity are easily configurable by the amount of sampling performed. Our approach, called Monte Carlo Quantization (MCQ), is linear in both time and space, with the resulting quantized, sparse networks showing minimal accuracy loss when compared to the original full-precision networks. Our method either outperforms or achieves competitive results on multiple benchmarks compared to previous quantization methods that do require additional training.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1905.12253 [cs.LG]
	(or arXiv:1905.12253v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1905.12253

Submission history

From: Alexander Keller [view email]
[v1] Wed, 29 May 2019 07:31:18 UTC (3,904 KB)
[v2] Tue, 7 Jan 2020 03:48:58 UTC (3,601 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2019-05

Change to browse by:

cs
stat
stat.ML

References & Citations

DBLP - CS Bibliography

listing | bibtex

Gonçalo Mordido
Matthijs Van Keirsbilck
Alexander Keller

export BibTeX citation

Computer Science > Machine Learning

Title:Instant Quantization of Neural Networks using Monte Carlo Methods

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Instant Quantization of Neural Networks using Monte Carlo Methods

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators