
While a vast amount of clustering algorithms of different types are available in the literature, the majority of existing algorithms depend on carefully tuned parameters to obtain satisfactory results. In this paper, we reduce the dependence on parameters on the basis of the dominant sets algorithm. The dominant sets algorithm is a parameter-independent clustering approach, which uses the pairwise data similarity matrix as input. If the data for clustering are in the form of feature vectors, it is necessary to measure the data similarity and build the similarity matrix. With the commonly used Gaussian kernel, the involved parameter is found to exert a significant influence on the clustering results. We study in depth why and how the dominant sets clustering results are influenced by the parameter and attribute the influence to the dominant set definition, which imposes a somewhat too strict constraint on internal similarity. A two-step clustering algorithm is then proposed to solve this problem. First, we transform the similarity matrix by histogram equalization before clustering, and this is shown to eliminate the influence of similarity parameter effectively. In the second step, we expand the clusters to maximize the ratio of internal similarity with respect to external similarity. Our algorithm is designed to achieve the balance between high internal similarity and low external similarity, thereby relieving the dependence on the similarity parameter. In experiments on ten publicly available data sets, our algorithm is shown to perform well in comparison with several other algorithms which benefit from carefully tuned parameters.
pattern classification, dominant set, Electrical engineering. Electronics. Nuclear engineering, fault diagnosis, Clustering, TK1-9971
pattern classification, dominant set, Electrical engineering. Electronics. Nuclear engineering, fault diagnosis, Clustering, TK1-9971
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 3 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
