On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)

Hu, Jerry Yao-Chieh; Wu, Weimin; Song, Zhao; Liu, Han

Statistics > Machine Learning

arXiv:2407.01079 (stat)

[Submitted on 1 Jul 2024 (v1), last revised 31 Oct 2024 (this version, v3)]

Title:On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)

Authors:Jerry Yao-Chieh Hu, Weimin Wu, Zhao Song, Han Liu

View PDF

Abstract:We investigate the statistical and computational limits of latent Diffusion Transformers (DiTs) under the low-dimensional linear latent space assumption. Statistically, we study the universal approximation and sample complexity of the DiTs score function, as well as the distribution recovery property of the initial data. Specifically, under mild data assumptions, we derive an approximation error bound for the score network of latent DiTs, which is sub-linear in the latent space dimension. Additionally, we derive the corresponding sample complexity bound and show that the data distribution generated from the estimated score function converges toward a proximate area of the original one. Computationally, we characterize the hardness of both forward inference and backward computation of latent DiTs, assuming the Strong Exponential Time Hypothesis (SETH). For forward inference, we identify efficient criteria for all possible latent DiTs inference algorithms and showcase our theory by pushing the efficiency toward almost-linear time inference. For backward computation, we leverage the low-rank structure within the gradient computation of DiTs training for possible algorithmic speedup. Specifically, we show that such speedup achieves almost-linear time latent DiTs training by casting the DiTs gradient as a series of chained low-rank approximations with bounded error. Under the low-dimensional assumption, we show that the statistical rates and the computational efficiency are all dominated by the dimension of the subspace, suggesting that latent DiTs have the potential to bypass the challenges associated with the high dimensionality of initial data.

Comments:	Accepted at NeurIPS 2024. v3 updated to camera-ready version with many typos fixed; v2 fixed typos, added Fig. 1 and added clarifications
Subjects:	Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2407.01079 [stat.ML]
	(or arXiv:2407.01079v3 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2407.01079

Submission history

From: Jerry Yao-Chieh Hu [view email]
[v1] Mon, 1 Jul 2024 08:34:40 UTC (72 KB)
[v2] Thu, 22 Aug 2024 06:25:19 UTC (336 KB)
[v3] Thu, 31 Oct 2024 16:59:13 UTC (85 KB)

Statistics > Machine Learning

Title:On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators