Downloads provided by UsageCounts
handle: 20.500.12358/25181
{"references": ["Angiulli, F., Pizzuti, C.: Fast Outlier detection in high dimensional\nspaces, In Proc. of the Sixth European Conference on the Principles of\nData Mining and Knowledge Discovery, pp. 15-26, 2002.", "Barbar\u251c\u00e1, D., Chen, P.: Using the fractal dimension to cluster datasets,\nIn: Proc. KDD, pp. 260-264, 2000.", "Barnett, V., Lewis, T.: Outliers in Statistical Data, John Wiley, 1994.", "Bay, S. D., and Schwabacher, M.: Mining Distance-Based Outliers in\nNear Linear Time with Randomization and a Simple Pruning Rule, Proc.\nof The Ninth ACM SIGKDD International Conference on Knowledge\nDiscovery and Data Mining, 2003.", "Blake C., Keogh E., Merz C. J.: UCI Repository of Machine Learning\nDatabases, http://www.ics.uci.edu/~mlearn/MLRepository.htm, 1998.", "Bolton, R. J., Hand, D. J.: Statistical fraud detection: A review (with\ndiscussion), Statistical Science, 17(3): pp. 235-255, 2002.", "Breunig, M., Kriegel, H., Ng, R., Sander, J.: LOF: Identifying densitybased\nlocal outliers, In: Proc. SIGMOD Conf, pp. 93-104, 2000.", "Eskin E., Arnold A., Prerau M., Portnoy L., Stolfo S.: A geometric\nframework for unsupervised anomaly detection: Detecting intrusions in\nunlabeled data, In Data Mining for Security Applications, 2002.", "Ester M., Kriegel H.-P., Sander J., Xu X.: A Density-Based Algorithm\nfor Discovering Clusters in Large Spatial Databases with Noise, Proc.\n2nd Int. Conf. on Knowledge Discovery and Data Mining (KDD'96),\nPortland, OR. pp. 226-231, 1996.\n[10] Han, J., Kamber, M.: Data Mining: Concepts and Techniques, San\nFrancisco, Morgan Kaufmann, 2001.\n[11] Hawkins, D.: Identification of Outliers, Chapman and Hall, 1980.\n[12] Hawkins, S., He, H. X., Williams, G. J., Baxter, R. A.: Outlier detection\nusing replicator neural networks, In Proc. of the Fifth Int. Conf. and\nData Warehousing and Knowledge Discovery (DaWaK02), 2002.\n[13] He, Z., Deng, S., Xu., X.: Outlier detection integrating semantic\nknowledge, In: Proc. of WAIM-02, pp. 126-131, 2002.\n[14] He, Z., Xu, X., Huang, J., Deng, S.: Mining Class Outliers: Concepts,\nAlgorithms and Applications in CRM, Expert Systems with Applications\n(ESWA'04), 27(4): pp. 681-697, 2004.\n[15] Jain, A., Murty, M., Flynn, P.: Data clustering: A review, ACM Comp,\nSurveys 31, 264-323, 1999.\n[16] Johnson, T., Kwok, I., Ng, R.: Fast computation of 2-dimensional depth\ncontours, In: Proc. KDD. pp. 224-228, 1998.\n[17] Knorr E. M., Ng. R. T.: Finding intensional knowledge of distancebased\noutliers, In Proc. of the 25th VLDB Conference, 1999.\n[18] Knorr, E., Ng, R., Tucakov, V.: Distance-based outliers: Algorithms and\napplications, VLDB Journal 8, pp. 237-253, 2000.\n[19] Knorr, E., Ng, R.: A unified notion of outliers: Properties and\ncomputation, In: Proc. KDD. pp. 219-222, 1997.\n[20] Knorr, E., Ng, R.: Finding intentional knowledge of distance-based\noutliers, In: Proc. VLDB. pp. 211-222, 1999.\n[21] Knorr, E.M., Ng, R.: Algorithms for mining distance-based outliers in\nlarge datasets, In: Proc. VLDB pp. 392-403, 1998.\n[22] Lane, T., Brodley, C. E.: Temporal sequence learning and data\nreduction for anomaly detection, ACM Transactions on Information and\nSystem Security, 2(3): pp. 295-331, 1999.\n[23] Michalski, R. S., Winston, P. H.: Variable Precision Logic, Artificial\nIntelligence Journal 29, Elsevier Science Publishers B.V. (North-\nHolland), pp. 121-146,1986.\n[24] Papadimitriou, S., Faloutsos C.: Cross-outlier detection, In: Proc. of\nSSTD-03, pp. 199-213, 2003.\n[25] Ramaswamy, S., Rastogi, R., Shim, K.: Efficient algorithms for mining\noutliers from large data sets, In Proc. of the ACM SIGMOD\nConference, pp. 427-438, 2000.\n[26] Rousseeuw, P., Leroy, A.: Robust Regression and Outlier Detection,\nJohn Wiley and Sons, 1987.\n[27] Rulequest Research, Gritbot, http://www.rulequest.com\n[28] Witten, I. H., Frank, E.: Data Mining: Practical Machine Learning Tools\nand Techniques, (Second Edition), San Francisco, Morgan Kaufmann,\n2005."]}
In large datasets, identifying exceptional or rare cases with respect to a group of similar cases is considered very significant problem. The traditional problem (Outlier Mining) is to find exception or rare cases in a dataset irrespective of the class label of these cases, they are considered rare events with respect to the whole dataset. In this research, we pose the problem that is Class Outliers Mining and a method to find out those outliers. The general definition of this problem is "given a set of observations with class labels, find those that arouse suspicions, taking into account the class labels". We introduce a novel definition of Outlier that is Class Outlier, and propose the Class Outlier Factor (COF) which measures the degree of being a Class Outlier for a data object. Our work includes a proposal of a new algorithm towards mining of the Class Outliers, presenting experimental results applied on various domains of real world datasets and finally a comparison study with other related methods is performed.
Outliers Mining., Distance-Based Approach, Class Outliers
Outliers Mining., Distance-Based Approach, Class Outliers
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 3 | |
| downloads | 6 |

Views provided by UsageCounts
Downloads provided by UsageCounts