Measuring the Validity of Clustering Validation Datasets

Name: Measuring the Validity of Clustering Validation Datasets
Keywords: FOS: Computer and information sciences, Computer Science - Machine Learning, Machine Learning (cs.LG)

Hyeon Jeon; Michaël Aupetit 0001; DongHwa Shin; Aeri Cho; Seokhyeon Park; Jinwook Seo

Found an issue? Give us feedback

arXiv.org e-Print Ar...arrow_drop_down

arXiv.org e-Print Archive

Preprint . 2025

Data sources: arXiv.org e-Print Archive

IEEE Transactions on Pattern Analysis and Machine Intelligence

Article . 2025 . Peer-reviewed

License: IEEE Copyright

Data sources: Crossref

https://dx.doi.org/10.48550/ar...

Article . 2025

License: CC BY NC ND

Data sources: Datacite

https://pubmed.ncbi.nlm.nih.go...

Article

Data sources: Europe PubMed Central

DBLP

Article

Data sources: DBLP

DBLP

Article

Data sources: DBLP

Measuring the Validity of Clustering Validation Datasets

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 01 Jun 2025Embargo end date: 01 Jan 2025Publisher:Institute of Electrical and Electronics Engineers (IEEE)Journal:IEEE Transactions on Pattern Analysis and Machine Intelligence, volume 47, pages 5,045-5,058 (issn: 0162-8828, eissn: 1939-3539,

Copyright policy )

Authors: Hyeon Jeon; Michaël Aupetit 0001; DongHwa Shin; Aeri Cho; Seokhyeon Park; Jinwook Seo;

doi: 10.1109/tpami.2025.3548011 , 10.48550/arxiv.2503.01097

pmid: 40036522

arXiv: 2503.01097

Measuring the Validity of Clustering Validation Datasets

- Summary
- Subjects
- Metrics

Abstract

Clustering techniques are often validated using benchmark datasets where class labels are used as ground-truth clusters. However, depending on the datasets, class labels may not align with the actual data clusters, and such misalignment hampers accurate validation. Therefore, it is essential to evaluate and compare datasets regarding their cluster-label matching (CLM), i.e., how well their class labels match actual clusters. Internal validation measures (IVMs), like Silhouette, can compare CLM over different labeling of the same dataset, but are not designed to do so across different datasets. We thus introduce Adjusted IVMs as fast and reliable methods to evaluate and compare CLM across datasets. We establish four axioms that require validation measures to be independent of data properties not related to cluster structure (e.g., dimensionality, dataset size). Then, we develop standardized protocols to convert any IVM to satisfy these axioms, and use these protocols to adjust six widely used IVMs. Quantitative experiments (1) verify the necessity and effectiveness of our protocols and (2) show that adjusted IVMs outperform the competitors, including standard IVMs, in accurately evaluating CLM both within and across datasets. We also show that the datasets can be filtered or improved using our method to form more reliable benchmarks for clustering validation.

IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

Related Organizations

Seoul National University
Korea (Republic of)
Kwangwoon University
Korea (Republic of)
Hamad bin Khalifa University
Qatar

Keywords

FOS: Computer and information sciences, Computer Science - Machine Learning, Machine Learning (cs.LG)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	2
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

2

Top 10%

Average

Green