Transformer Block Coupling and its Correlation with Generalization in LLMs

Aubry, Murdock; Meng, Haoming; Sugolov, Anton; Papyan, Vardan

Computer Science > Machine Learning

arXiv:2407.07810 (cs)

[Submitted on 10 Jul 2024 (v1), last revised 5 Mar 2025 (this version, v5)]

Title:Transformer Block Coupling and its Correlation with Generalization in LLMs

Authors:Murdock Aubry, Haoming Meng, Anton Sugolov, Vardan Papyan

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) have made significant strides in natural language processing, and a precise understanding of the internal mechanisms driving their success is essential. In this work, we analyze the trajectories of token embeddings as they pass through transformer blocks, linearizing the system along these trajectories through their Jacobian matrices. By examining the relationships between these block Jacobians, we uncover the phenomenon of \textbf{transformer block coupling} in a multitude of LLMs, characterized by the coupling of their top singular vectors across tokens and depth. Our findings reveal that coupling \textit{positively correlates} with model performance, and that this relationship is stronger than with other hyperparameters such as parameter count, model depth, and embedding dimension. We further investigate how these properties emerge during training, observing a progressive development of coupling, increased linearity, and layer-wise exponential growth in token trajectories. Additionally, experiments with Vision Transformers (ViTs) corroborate the emergence of coupling and its relationship with generalization, reinforcing our findings in LLMs. Collectively, these insights offer a novel perspective on token interactions in transformers, opening new directions for studying their mechanisms as well as improving training and generalization.

Comments:	Published as a conference paper at the International Conference on Learning Representations (ICLR 2025)
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2407.07810 [cs.LG]
	(or arXiv:2407.07810v5 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2407.07810

Submission history

From: Haoming Meng [view email]
[v1] Wed, 10 Jul 2024 16:30:27 UTC (43,199 KB)
[v2] Mon, 14 Oct 2024 04:29:05 UTC (40,676 KB)
[v3] Sun, 22 Dec 2024 06:54:57 UTC (43,977 KB)
[v4] Mon, 3 Mar 2025 19:51:59 UTC (43,690 KB)
[v5] Wed, 5 Mar 2025 04:47:05 UTC (43,690 KB)

Computer Science > Machine Learning

Title:Transformer Block Coupling and its Correlation with Generalization in LLMs

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Transformer Block Coupling and its Correlation with Generalization in LLMs

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators