Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix

Hayase, Tomohiro; Collins, Benoît; Karakida, Ryo

Statistics > Machine Learning

arXiv:2510.06685 (stat)

[Submitted on 8 Oct 2025]

Title:Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix

Authors:Tomohiro Hayase, Benoît Collins, Ryo Karakida

View PDF

Abstract:Self-attention layers have become fundamental building blocks of modern deep neural networks, yet their theoretical understanding remains limited, particularly from the perspective of random matrix theory. In this work, we provide a rigorous analysis of the singular value spectrum of the attention matrix and establish the first Gaussian equivalence result for attention. In a natural regime where the inverse temperature remains of constant order, we show that the singular value distribution of the attention matrix is asymptotically characterized by a tractable linear model. We further demonstrate that the distribution of squared singular values deviates from the Marchenko-Pastur law, which has been believed in previous work. Our proof relies on two key ingredients: precise control of fluctuations in the normalization term and a refined linearization that leverages favorable Taylor expansions of the exponential. This analysis also identifies a threshold for linearization and elucidates why attention, despite not being an entrywise operation, admits a rigorous Gaussian equivalence in this regime.

Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG); Probability (math.PR)
Cite as:	arXiv:2510.06685 [stat.ML]
	(or arXiv:2510.06685v1 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2510.06685

Submission history

From: Tomohiro Hayase [view email]
[v1] Wed, 8 Oct 2025 06:13:42 UTC (142 KB)

Statistics > Machine Learning

Title:Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Gaussian Equivalence for Self-Attention: Asymptotic Spectral Analysis of Attention Matrix

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators