Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Article . 2007
License: CC BY
Data sources: Datacite
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Article . 2007
License: CC BY
Data sources: ZENODO
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Article . 2007
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

Categorical Clustering By Converting Associated Information

Authors: Dongmin Cai; Yau, Stephen S-T;

Categorical Clustering By Converting Associated Information

Abstract

{"references": ["C. C. Aggarwal, A human-computer interactive method for projected\nclustering, IEEE Transactions on Knowledge and Data Engineering,\n16(4), 448-460, 2004.", "M. Ankerst, M. Breunig, H.-P. Kriegel, and J. Sander. OPTICS:\nOrdering points to identify the clustering structure. In Proc. 1999 ACMSIGMOD\nInt. Conf. Management of Data (SIGMOD'99), pages 49{60,\nPhiladelphia, PA, June 1999.", "M.R. Anderberg, Cluster analysis for applications, Academic Press,\n1973.", "D. Barbara, Y. Li, J. Couto, COOLCAT: An entropy-based algorithm for\ncategorical clustering. In: CIKM Conference. McLean, VA, 2002.", "C.L. Blake and C.J. Merz, UCI repository of machine learning\ndatabases, 1998. http://www.ics.uci.edu/~mlearn/MLRepository.html", "D. Cristofor and D. A. Simovici, An information-theoretical approach to\nclustering categorical databases using genetic algorithms. In Proceedings\nof the Workshop on Clustering High-Dimensional Data and Its\nApplications (SIAM ICDM), pages 37-46, Washington, 2002.", "Richard O. Duda and Peter E. Hard, Pattern classification and scene\nanalysi. A wiley-Interscience Publication, New York, 1973.", "M. Ester, H.-P. Kriegel, J. Sander, and X. Xu. A density-based algorithm\nfor discovering clusters in large spatial databases. In Proc. 1996 Int.\nConf. Knowledge Discovery and Data Mining (KDD'96), pages\n226{231, Portland, Oregon, Aug. 1996.", "D. Fisher, Improving inference through conceptual clustering. In Proc.\n1987 National Conference Artificial Intelligence (AAAI-87), pages 461-\n465, Seattle, WA, July 1987.\n[10] K.C. Gowda and E. Diday, Symbolic clustering using a new dissimilarity\nmeasure. Pattern Recognition, 24(6): 567-578, 1991.\n[11] V. Ganti, J. Gehrke, and R. Ramakrishnan. CACTUS: Clustering\ncategorical data using summaries. In ACM SIGKDD Int-l Conference on\nKnowledge discovery in Databases, 1999.\n[12] David Gibson, Jon Kleiberg, Prabhakar Raghavan: Clustering\ncategorical data: an approach based on dynamic systems\". Proc. 1998\nInt. Conf. On Very Large Databases, pp. 311-323, New York, August\n1998.\n[13] J.C. Gower, A general coefficient of similarity and some of its\nproperties. BioMetrics, 27: 857-874, 1971.\n[14] Sudipto Guha, Rajeev Rastogi, Kyuseok Shim, ROCK: A robust\nclustering algorithm for categorical attributes. ICDE 1999: 512-521.\n[15] A. Hinneburg and D. A. Keim. An efficient approach to clustering in\nlarge multimedia databases with noise. In Proc. 1998 Int. Conf.\nKnowledge Discovery and Data Mining (KDD'98), pages 58-65, New\nYork, NY, Aug. 1998.\n[16] J. Han and M. Kamber, Data mining: concepts and techniques, Morgan\nKaufmann publishers, 2001.\n[17] Z. Huang, Extensions to the k-means algorithm for clustering large data\nsets with categorical values, Data Mining and Knowledge Discovery,\nvol. 2, no. 3, pp 283-304, 1998.\n[18] A.K. Jain and R.C. Dubes, Algorithms for clustering data, Rentice Hall,\n1988.\n[19] L. Kaufman and P.J. Rousseeuw, Finding groups in data - An\nIntroduction to Cluster Analysis in Knowledge, 1990.\n[20] Lioyd. Learning square quantization in PCM. (published in IEEE Trans.\nInformation Theory), 28:128-137, 1982), Technical Report, Bell Labs,\n1957.\n[21] Tao Li, Sheng Ma, Mitsunori Ogihara, Entropy-based criterion in\ncategorical clustering. In Proceedings of The 2004, IEEE International\nConference on Machine Learning (ICML 2004), pages 536-543.\n[22] J. MacQueen. Some methods for classi\u252c\u00bbcation and analysis of\nmultivariate observations. Proc. 5th Berkeley Symp. Math. Statist, Prob.,\n1:281-297, 1967.\n[23] R.S. Michalski and R.E. Stephen, Automated construction of\nclassification: conceptual clustering versus numerical taxonomy. IEEE\nTransactions on Pattern Analysis and Machine Intelligence, 5(4): 396-\n410, 1983.\n[24] J.R. Quinlan, Induction of decision trees, Machine Learning, vol. 1, no.\n1, pp. 81-106, 1986.\n[25] J.R. Quinlan, C4.5: Programs for machine learning. Morgan Kaufmann,\n1993.\n[26] H. Ralambondrainy, A conceptual version of the k-means algorithm.\nPattern Recognition Letters, 16:1147-1157, 1995.\n[27] Claude. E. Shannon, A mathematical theory of communication, Bell\nSystem Technical Journal, vol.27, pp. 379-423 and 623-656, July and\nOctober, 1948.\n[28] M. Steinbach, G. Karypis, and V. Kumar, A comparison of document\nclustering techniques, In KDD workshop on Text Mining, 2000.\n[29] L. Talavera and J. B\u00e9jar, Intergrating declarative knowledge in\nhierarchical clustering tasks. Proceedings of the International\nSymposium on Intelligent Data Analysis, pp. 211-222, Amsterdam, The\nNetherlands: Springer-Verlag, 1999.\n[30] Y. Zhang, A. Fu, C. Cai, and P. Heng, Clustering categorical data, In\nProc. 2000 IEEE Int. Conf. Data Engineering, San Deigo, USA, March\n2000."]}

Lacking an inherent "natural" dissimilarity measure between objects in categorical dataset presents special difficulties in clustering analysis. However, each categorical attributes from a given dataset provides natural probability and information in the sense of Shannon. In this paper, we proposed a novel method which heuristically converts categorical attributes to numerical values by exploiting such associated information. We conduct an experimental study with real-life categorical dataset. The experiment demonstrates the effectiveness of our approach.

Keywords

Information, Categorical, Clustering, Converting

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
    OpenAIRE UsageCounts
    Usage byUsageCounts
    visibility views 2
    download downloads 6
  • 2
    views
    6
    downloads
    Powered byOpenAIRE UsageCounts
Powered by OpenAIRE graph
Found an issue? Give us feedback
visibility
download
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
views
OpenAIRE UsageCountsViews provided by UsageCounts
downloads
OpenAIRE UsageCountsDownloads provided by UsageCounts
0
Average
Average
Average
2
6
Green