Sparse Expansion and Neuronal Disentanglement

Sawmya, Shashata; Kong, Linghao; Markov, Ilia; Alistarh, Dan; Shavit, Nir

Computer Science > Machine Learning

arXiv:2405.15756v1 (cs)

[Submitted on 24 May 2024 (this version), latest version 26 Feb 2025 (v4)]

Title:Sparse Expansion and Neuronal Disentanglement

Authors:Shashata Sawmya, Linghao Kong, Ilia Markov, Dan Alistarh, Nir Shavit

View PDF HTML (experimental)

Abstract:We show how to improve the inference efficiency of an LLM by expanding it into a mixture of sparse experts, where each expert is a copy of the original weights, one-shot pruned for a specific cluster of input values. We call this approach $\textit{Sparse Expansion}$. We show that, for models such as Llama 2 70B, as we increase the number of sparse experts, Sparse Expansion outperforms all other one-shot sparsification approaches for the same inference FLOP budget per token, and that this gap grows as sparsity increases, leading to inference speedups.
But why? To answer this, we provide strong evidence that the mixture of sparse experts is effectively $\textit{disentangling}$ the input-output relationship of every individual neuron across clusters of inputs. Specifically, sparse experts approximate the dense neuron output distribution with fewer weights by decomposing the distribution into a collection of simpler ones, each with a separate sparse dot product covering it. Interestingly, we show that the Wasserstein distance between a neuron's output distribution and a Gaussian distribution is an indicator of its entanglement level and contribution to the accuracy of the model. Every layer of an LLM has a fraction of highly entangled Wasserstein neurons, and model performance suffers more when these are sparsified as opposed to others.

Comments:	9 pages, 8 figures, Submitted to NeurIPS 2024 main track
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2405.15756 [cs.LG]
	(or arXiv:2405.15756v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2405.15756

Submission history

From: Shashata Sawmya [view email]
[v1] Fri, 24 May 2024 17:51:39 UTC (5,263 KB)
[v2] Mon, 24 Jun 2024 22:14:42 UTC (5,263 KB)
[v3] Mon, 17 Feb 2025 01:06:24 UTC (11,484 KB)
[v4] Wed, 26 Feb 2025 17:32:10 UTC (11,484 KB)

Computer Science > Machine Learning

Title:Sparse Expansion and Neuronal Disentanglement

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Sparse Expansion and Neuronal Disentanglement

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators