The Asymptotic Behavior of Attention in Transformers

Abella, Álvaro Rodríguez; Silvestre, João Pedro; Tabuada, Paulo

Computer Science > Artificial Intelligence

arXiv:2412.02682 (cs)

[Submitted on 3 Dec 2024 (v1), last revised 25 Sep 2025 (this version, v2)]

Title:The Asymptotic Behavior of Attention in Transformers

Authors:Álvaro Rodríguez Abella, João Pedro Silvestre, Paulo Tabuada

View PDF HTML (experimental)

Abstract:The transformer architecture has become the foundation of modern Large Language Models (LLMs), yet its theoretical properties are still not well understood. As with classic neural networks, a common approach to improve these models is to increase their size and depth. However, such strategies may be suboptimal, as several works have shown that adding more layers yields increasingly diminishing returns. More importantly, prior studies have shown that increasing depth may lead to model collapse, i.e., all the tokens converge to a single cluster, undermining the ability of LLMs to generate diverse outputs. Building on differential equation models for the transformer dynamics, we prove that all the tokens in a transformer asymptotically converge to a cluster as depth increases. At the technical level we leverage tools from control theory, including consensus dynamics on manifolds and input-to-state stability (ISS). We then extend our analysis to autoregressive models, exploiting their structure to further generalize the theoretical guarantees.

Subjects:	Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Systems and Control (eess.SY); Dynamical Systems (math.DS); Optimization and Control (math.OC)
Cite as:	arXiv:2412.02682 [cs.AI]
	(or arXiv:2412.02682v2 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2412.02682

Submission history

From: João Pedro Silvestre [view email]
[v1] Tue, 3 Dec 2024 18:54:49 UTC (3,197 KB)
[v2] Thu, 25 Sep 2025 00:09:18 UTC (1,331 KB)

Computer Science > Artificial Intelligence

Title:The Asymptotic Behavior of Attention in Transformers

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:The Asymptotic Behavior of Attention in Transformers

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators