
Clustering in high-dimensional spaces is a difficult problem which is recurrent in many domains, for example in image analysis. The difficulty is due to the fact that high-dimensional data usually live in different low-dimensional subspaces hidden in the original space. This paper presents a family of Gaussian mixture models designed for high-dimensional data which combine the ideas of dimension reduction and parsimonious modeling. These models give rise to a clustering method based on the Expectation-Maximization algorithm which is called High-Dimensional Data Clustering (HDDC). In order to correctly fit the data, HDDC estimates the specific subspace and the intrinsic dimension of each group. Our experiments on artificial and real datasets show that HDDC outperforms existing methods for clustering high-dimensional data
ACM-G3-Multivariate statistics, Classification and discrimination; cluster analysis (statistical aspects), dimension reduction, [STAT.TH] Statistics [stat]/Statistics Theory [stat.TH], [INFO.INFO-CV]Computer Science [cs]/Computer Vision and Pattern Recognition [cs.CV], 500, model-based clustering, Mathematics - Statistics Theory, [STAT.TH]Statistics [stat]/Statistics Theory [stat.TH], Statistics Theory (math.ST), 510, subspace clustering, Model-based clustering, [INFO.INFO-CV] Computer Science [cs]/Computer Vision and Pattern Recognition [cs.CV], high-dimensional data, [MATH.MATH-ST]Mathematics [math]/Statistics [math.ST], subspace selection, FOS: Mathematics, Gaussian mixture models, Computational methods for problems pertaining to statistics, [MATH.MATH-ST] Mathematics [math]/Statistics [math.ST], parsimonious models
ACM-G3-Multivariate statistics, Classification and discrimination; cluster analysis (statistical aspects), dimension reduction, [STAT.TH] Statistics [stat]/Statistics Theory [stat.TH], [INFO.INFO-CV]Computer Science [cs]/Computer Vision and Pattern Recognition [cs.CV], 500, model-based clustering, Mathematics - Statistics Theory, [STAT.TH]Statistics [stat]/Statistics Theory [stat.TH], Statistics Theory (math.ST), 510, subspace clustering, Model-based clustering, [INFO.INFO-CV] Computer Science [cs]/Computer Vision and Pattern Recognition [cs.CV], high-dimensional data, [MATH.MATH-ST]Mathematics [math]/Statistics [math.ST], subspace selection, FOS: Mathematics, Gaussian mixture models, Computational methods for problems pertaining to statistics, [MATH.MATH-ST] Mathematics [math]/Statistics [math.ST], parsimonious models
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 286 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 1% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 1% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
