Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior

Han, Fuqun; Osher, Stanley; Li, Wuchen

Computer Science > Machine Learning

arXiv:2510.16356 (cs)

[Submitted on 18 Oct 2025]

Title:Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior

Authors:Fuqun Han, Stanley Osher, Wuchen Li

View PDF HTML (experimental)

Abstract:In this work, we propose a sparse transformer architecture that incorporates prior information about the underlying data distribution directly into the transformer structure of the neural network. The design of the model is motivated by a special optimal transport problem, namely the regularized Wasserstein proximal operator, which admits a closed-form solution and turns out to be a special representation of transformer architectures. Compared with classical flow-based models, the proposed approach improves the convexity properties of the optimization problem and promotes sparsity in the generated samples. Through both theoretical analysis and numerical experiments, including applications in generative modeling and Bayesian inverse problems, we demonstrate that the sparse transformer achieves higher accuracy and faster convergence to the target distribution than classical neural ODE-based methods.

Subjects:	Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
Cite as:	arXiv:2510.16356 [cs.LG]
	(or arXiv:2510.16356v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2510.16356

Submission history

From: Fuqun Han [view email]
[v1] Sat, 18 Oct 2025 05:26:13 UTC (3,420 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2025-10

Change to browse by:

cs
math
math.OC
stat
stat.ML

References & Citations

export BibTeX citation

Computer Science > Machine Learning

Title:Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Sparse Transformer Architectures via Regularized Wasserstein Proximal Operator with $L_1$ Prior

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators