Validation of cluster analysis results on validation data: A systematic framework

Ullmann, Theresa; Hennig, Christian; Boulesteix, Anne-Laure

Statistics > Methodology

arXiv:2103.01281 (stat)

[Submitted on 1 Mar 2021 (v1), last revised 10 Jan 2022 (this version, v2)]

Title:Validation of cluster analysis results on validation data: A systematic framework

Authors:Theresa Ullmann, Christian Hennig, Anne-Laure Boulesteix

View PDF

Abstract:Cluster analysis refers to a wide range of data analytic techniques for class discovery and is popular in many application fields. To judge the quality of a clustering result, different cluster validation procedures have been proposed in the literature. While there is extensive work on classical validation techniques, such as internal and external validation, less attention has been given to validating and replicating a clustering result using a validation dataset. Such a dataset may be part of the original dataset, which is separated before analysis begins, or it could be an independently collected dataset. We present a systematic structured framework for validating clustering results on validation data that includes most existing validation approaches. In particular, we review classical validation techniques such as internal and external validation, stability analysis, hypothesis testing, and visual validation, and show how they can be interpreted in terms of our framework. We precisely define and formalise different types of validation of clustering results on a validation dataset and explain how each type can be implemented in practice. Furthermore, we give examples of how clustering studies from the applied literature that used a validation dataset can be classified into the framework.

Comments:	32 pages, 1 figure
Subjects:	Methodology (stat.ME)
Cite as:	arXiv:2103.01281 [stat.ME]
	(or arXiv:2103.01281v2 [stat.ME] for this version)
	https://doi.org/10.48550/arXiv.2103.01281

Submission history

From: Theresa Ullmann [view email]
[v1] Mon, 1 Mar 2021 19:53:59 UTC (63 KB)
[v2] Mon, 10 Jan 2022 14:30:27 UTC (64 KB)

Statistics > Methodology

Title:Validation of cluster analysis results on validation data: A systematic framework

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Methodology

Title:Validation of cluster analysis results on validation data: A systematic framework

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators