The Cluster Structure Function

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 01 Sep 2023Embargo end date: 01 Jan 2022 Netherlands Publisher:Institute of Electrical and Electronics Engineers (IEEE)Journal:IEEE Transactions on Pattern Analysis and Machine Intelligence, volume 45, pages 11,309-11,320 (issn: 0162-8828, eissn: 1939-3539,

Copyright policy )Funded by:NIH | Quantifying Changes in Ne...

Authors: Andrew R. Cohen; Paul M. B. Vitányi;

doi: 10.1109/tpami.2023.3264690 , 10.48550/arxiv.2201.01222

pmid: 37018105

pmc: PMC10525042

arXiv: 2201.01222

The Cluster Structure Function

- Summary
- Subjects
- External Databases
  (1)
- Metrics

Abstract

For each partition of a data set into a given number of parts there is a partition such that every part is as much as possible a good model (an "algorithmic sufficient statistic") for the data in that part. Since this can be done for every number between one and the number of data, the result is a function, the cluster structure function. It maps the number of parts of a partition to values related to the deficiencies of being good models by the parts. Such a function starts with a value at least zero for no partition of the data set and descents to zero for the partition of the data set into singleton parts. The optimal clustering is the one chosen to minimize the cluster structure function. The theory behind the method is expressed in algorithmic information theory (Kolmogorov complexity). In practice the Kolmogorov complexities involved are approximated by a concrete compressor. We give examples using real data sets: the MNIST handwritten digits and the segmentation of real cells as used in stem cell research.

Country

Netherlands

Related Organizations

Centrum Wiskunde & Informatica
Netherlands
Drexel University
United States
University of Amsterdam
Netherlands
Dutch Research Council
Netherlands

Keywords

FOS: Computer and information sciences, Computer Science - Machine Learning, Pattern, Computer Vision and Pattern Recognition (cs.CV), Computer Science - Computer Vision and Pattern Recognition, Kolmogorov complexity, Classification, Similarity, Article, Algorithmic sufficient statistic, Machine Learning (cs.LG), Cluster, recognition, Data mining

1the

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average