Downloads provided by UsageCounts
{"references": ["C. C. Aggarwal, A human-computer interactive method for projected\nclustering, IEEE Transactions on Knowledge and Data Engineering,\n16(4), 448-460, 2004.", "M. Ankerst, M. Breunig, H.-P. Kriegel, and J. Sander. OPTICS:\nOrdering points to identify the clustering structure. In Proc. 1999 ACMSIGMOD\nInt. Conf. Management of Data (SIGMOD'99), pages 49{60,\nPhiladelphia, PA, June 1999.", "M.R. Anderberg, Cluster analysis for applications, Academic Press,\n1973.", "D. Barbara, Y. Li, J. Couto, COOLCAT: An entropy-based algorithm for\ncategorical clustering. In: CIKM Conference. McLean, VA, 2002.", "C.L. Blake and C.J. Merz, UCI repository of machine learning\ndatabases, 1998. http://www.ics.uci.edu/~mlearn/MLRepository.html", "D. Cristofor and D. A. Simovici, An information-theoretical approach to\nclustering categorical databases using genetic algorithms. In Proceedings\nof the Workshop on Clustering High-Dimensional Data and Its\nApplications (SIAM ICDM), pages 37-46, Washington, 2002.", "Richard O. Duda and Peter E. Hard, Pattern classification and scene\nanalysi. A wiley-Interscience Publication, New York, 1973.", "M. Ester, H.-P. Kriegel, J. Sander, and X. Xu. A density-based algorithm\nfor discovering clusters in large spatial databases. In Proc. 1996 Int.\nConf. Knowledge Discovery and Data Mining (KDD'96), pages\n226{231, Portland, Oregon, Aug. 1996.", "D. Fisher, Improving inference through conceptual clustering. In Proc.\n1987 National Conference Artificial Intelligence (AAAI-87), pages 461-\n465, Seattle, WA, July 1987.\n[10] K.C. Gowda and E. Diday, Symbolic clustering using a new dissimilarity\nmeasure. Pattern Recognition, 24(6): 567-578, 1991.\n[11] V. Ganti, J. Gehrke, and R. Ramakrishnan. CACTUS: Clustering\ncategorical data using summaries. In ACM SIGKDD Int-l Conference on\nKnowledge discovery in Databases, 1999.\n[12] David Gibson, Jon Kleiberg, Prabhakar Raghavan: Clustering\ncategorical data: an approach based on dynamic systems\". Proc. 1998\nInt. Conf. On Very Large Databases, pp. 311-323, New York, August\n1998.\n[13] J.C. Gower, A general coefficient of similarity and some of its\nproperties. BioMetrics, 27: 857-874, 1971.\n[14] Sudipto Guha, Rajeev Rastogi, Kyuseok Shim, ROCK: A robust\nclustering algorithm for categorical attributes. ICDE 1999: 512-521.\n[15] A. Hinneburg and D. A. Keim. An efficient approach to clustering in\nlarge multimedia databases with noise. In Proc. 1998 Int. Conf.\nKnowledge Discovery and Data Mining (KDD'98), pages 58-65, New\nYork, NY, Aug. 1998.\n[16] J. Han and M. Kamber, Data mining: concepts and techniques, Morgan\nKaufmann publishers, 2001.\n[17] Z. Huang, Extensions to the k-means algorithm for clustering large data\nsets with categorical values, Data Mining and Knowledge Discovery,\nvol. 2, no. 3, pp 283-304, 1998.\n[18] A.K. Jain and R.C. Dubes, Algorithms for clustering data, Rentice Hall,\n1988.\n[19] L. Kaufman and P.J. Rousseeuw, Finding groups in data - An\nIntroduction to Cluster Analysis in Knowledge, 1990.\n[20] Lioyd. Learning square quantization in PCM. (published in IEEE Trans.\nInformation Theory), 28:128-137, 1982), Technical Report, Bell Labs,\n1957.\n[21] Tao Li, Sheng Ma, Mitsunori Ogihara, Entropy-based criterion in\ncategorical clustering. In Proceedings of The 2004, IEEE International\nConference on Machine Learning (ICML 2004), pages 536-543.\n[22] J. MacQueen. Some methods for classi\u252c\u00bbcation and analysis of\nmultivariate observations. Proc. 5th Berkeley Symp. Math. Statist, Prob.,\n1:281-297, 1967.\n[23] R.S. Michalski and R.E. Stephen, Automated construction of\nclassification: conceptual clustering versus numerical taxonomy. IEEE\nTransactions on Pattern Analysis and Machine Intelligence, 5(4): 396-\n410, 1983.\n[24] J.R. Quinlan, Induction of decision trees, Machine Learning, vol. 1, no.\n1, pp. 81-106, 1986.\n[25] J.R. Quinlan, C4.5: Programs for machine learning. Morgan Kaufmann,\n1993.\n[26] H. Ralambondrainy, A conceptual version of the k-means algorithm.\nPattern Recognition Letters, 16:1147-1157, 1995.\n[27] Claude. E. Shannon, A mathematical theory of communication, Bell\nSystem Technical Journal, vol.27, pp. 379-423 and 623-656, July and\nOctober, 1948.\n[28] M. Steinbach, G. Karypis, and V. Kumar, A comparison of document\nclustering techniques, In KDD workshop on Text Mining, 2000.\n[29] L. Talavera and J. B\u00e9jar, Intergrating declarative knowledge in\nhierarchical clustering tasks. Proceedings of the International\nSymposium on Intelligent Data Analysis, pp. 211-222, Amsterdam, The\nNetherlands: Springer-Verlag, 1999.\n[30] Y. Zhang, A. Fu, C. Cai, and P. Heng, Clustering categorical data, In\nProc. 2000 IEEE Int. Conf. Data Engineering, San Deigo, USA, March\n2000."]}
Lacking an inherent "natural" dissimilarity measure between objects in categorical dataset presents special difficulties in clustering analysis. However, each categorical attributes from a given dataset provides natural probability and information in the sense of Shannon. In this paper, we proposed a novel method which heuristically converts categorical attributes to numerical values by exploiting such associated information. We conduct an experimental study with real-life categorical dataset. The experiment demonstrates the effectiveness of our approach.
Information, Categorical, Clustering, Converting
Information, Categorical, Clustering, Converting
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 2 | |
| downloads | 6 |

Views provided by UsageCounts
Downloads provided by UsageCounts