Scalable Clustering: Large Scale Unsupervised Learning of Gaussian Mixture Models with Outliers

Zhou, Yijia; Gallivan, Kyle A.; Barbu, Adrian

doi:10.1080/10618600.2024.2414889

Statistics > Machine Learning

arXiv:2302.14599 (stat)

[Submitted on 28 Feb 2023]

Title:Scalable Clustering: Large Scale Unsupervised Learning of Gaussian Mixture Models with Outliers

Authors:Yijia Zhou, Kyle A. Gallivan, Adrian Barbu

View PDF

Abstract:Clustering is a widely used technique with a long and rich history in a variety of areas. However, most existing algorithms do not scale well to large datasets, or are missing theoretical guarantees of convergence. This paper introduces a provably robust clustering algorithm based on loss minimization that performs well on Gaussian mixture models with outliers. It provides theoretical guarantees that the algorithm obtains high accuracy with high probability under certain assumptions. Moreover, it can also be used as an initialization strategy for $k$-means clustering. Experiments on real-world large-scale datasets demonstrate the effectiveness of the algorithm when clustering a large number of clusters, and a $k$-means algorithm initialized by the algorithm outperforms many of the classic clustering methods in both speed and accuracy, while scaling well to large datasets such as ImageNet.

Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as:	arXiv:2302.14599 [stat.ML]
	(or arXiv:2302.14599v1 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2302.14599
Related DOI:	https://doi.org/10.1080/10618600.2024.2414889

Submission history

From: Yijia Zhou [view email]
[v1] Tue, 28 Feb 2023 14:39:18 UTC (820 KB)

Statistics > Machine Learning

Title:Scalable Clustering: Large Scale Unsupervised Learning of Gaussian Mixture Models with Outliers

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Scalable Clustering: Large Scale Unsupervised Learning of Gaussian Mixture Models with Outliers

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators