Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Article . 2017
License: CC BY
Data sources: Datacite
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Article . 2017
License: CC BY
Data sources: Datacite
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Article . 2017
License: CC BY
Data sources: ZENODO
versions View all 2 versions
addClaim

Fcnn-Mr: A Parallel Instance Selection Method Based On Fast Condensed Nearest Neighbor Rule

Authors: Lu Si; Jie Yu; Shasha Li; Jun Ma; Lei Luo; Qingbo Wu; Yongqi Ma; +1 Authors

Fcnn-Mr: A Parallel Instance Selection Method Based On Fast Condensed Nearest Neighbor Rule

Abstract

{"references": ["S. Garc\u0142a, J. Luengo, and F. Herrera, \"Data preprocessing in data\nmining,\" Computer Science, vol. 72, 2015.", "B. J. Han, \"Data mining. concepts and techniques. 3rd ed,\" Data Mining\nConcepts Models Methods & Algorithms Second Edition, vol. 5, no. 4,\npp. 1 \u2013 18, 2000.", "T. Cover and P. Hart, \"Nearest neighbor pattern classification,\" IEEE\nTransactions on Information Theory, vol. 13, no. 1, pp. 21\u201327, 1967.", "D. R. Wilson and T. R. Martinez, \"Reduction techniques for\ninstance-based learning algorithms,\" Machine Learning, vol. 38, no. 3,\npp. 257\u2013286, 2000.", "M. Kudo and J. Sklansky, \"Comparison of algorithms that select features\nfor pattern classifiers,\" Pattern Recognition, vol. 33, no. 1, pp. 25\u201341,\n2000.", "H. Liu and H. Motoda, \"Feature extraction construction and selection:\nA data mining perspective,\" Springer International, vol. 94, no. 448, p.\n014004, 1999.", "I. Triguero, J. Derrac, S. Garcia, and F. Herrera, \"A taxonomy\nand experimental study on prototype generation for nearest neighbor\nclassification,\" Systems Man & Cybernetics Part C Applications &\nReviews IEEE Transactions on, vol. 42, no. 1, pp. 86\u2013100, 2012.", "J. Hamidzadeh, R. Monsefi, and H. S. Yazdi, \"Irahc: Instance reduction\nalgorithm using hyperrectangle,\" Pattern Recognition, vol. 48, no. 5, pp.\n1878\u20131889, 2015.", "Haro-Garc, A. Aida, Garc, and N. A-Pedrajas, \"A divide-and-conquer\nrecursive approach for scaling up instance selection algorithms,\" Data\nMining and Knowledge Discovery, vol. 18, no. 3, pp. 392\u2013418, 2009.\n[10] I. Triguero, D. Peralta, J. Bacardit, S. Garc\u0142a, and F. Herrera, \"Mrpr: A\nmapreduce solution for prototype reduction in big data classification,\"\nNeurocomputing, vol. 150, no. 150, p. 331C345, 2015.\n[11] J. Zhai, X. Wang, and X. Pang, \"Voting-based instance selection\nfrom large data sets with mapreduce and random weight networks,\"\nInformation Sciences, vol. 367, pp. 1066\u20131077, 2016.\n[12] H. Liu and H. Motoda, \"On issues of instance selection.\" Data Mining\nand Knowledge Discovery, vol. 6, no. 2, pp. 115\u2013130, 2002.\n[13] J. R. Cano, F. Herrera, and M. Lozano, \"Stratification for scaling up\nevolutionary prototype selection,\" Pattern Recognition Letters, vol. 26,\nno. 7, pp. 953\u2013963, 2005.\n[14] J. Dean and S. Ghemawat, \"Mapreduce: Simplified data processing\non large clusters.\" in Conference on Symposium on Opearting Systems\nDesign & Implementation, 2004, pp. 107\u2013113.\n[15] F. Angiulli, \"Fast condensed nearest neighbor rule,\" in International\nConference, 2005, pp. 25\u201332.\n[16] Angiulli, \"Fast nearest neighbor condensation for large data sets\nclassification,\" IEEE Transactions on Knowledge & Data Engineering,\nvol. 19, no. 11, pp. 1450\u20131464, 2007.\n[17] B. P. E. Hart, \"The condensed nearest neighbor rule,\" in IEEE Trans.\nInformation Theory, 1968.\n[18] C. H. Chou, B. H. Kuo, and F. Chang, \"The generalized condensed\nnearest neighbor rule as a data reduction method,\" vol. 2, pp. 556\u2013559,\n2006.\n[19] G. W. Gates, \"The reduced nearest neighbor rule,\" IEEE Transactions\non Information Theory, vol. 18, no. 3, pp. 431 \u2013 433, 1972.\n[20] D. L. Wilson, \"Asymptotic properties of nearest neighbor rules using\nedited data,\" IEEE Transactions on Systems Man & Cybernetics, vol. 2,\nno. 3, pp. 408\u2013421, 1972.\n[21] W. C. Lin, C. F. Tsai, S. W. Ke, C. W. Hung, and W. Eberle, \"Learning\nto detect representative data for large scale instance selection,\" Journal\nof Systems & Software, vol. 106, no. C, pp. 1\u20138, 2015.\n[22] A. Onan, \"A fuzzy-rough nearest neighbor classifier combined with\nconsistency-based subset evaluation and instance selection for automated\ndiagnosis of breast cancer,\" Expert Systems with Applications, vol. 42,\nno. 20, pp. 6844\u20136852, 2015.\n[23] J. A. Olvera-Lpez, J. A. Carrasco-Ochoa, and J. F. Mart\u0142nez-Trinidad,\n\"A new fast prototype selection method based on clustering,\" Pattern\nAnalysis and Applications, vol. 13, no. 2, pp. 131\u2013141, 2010.\n[24] J. R. Cano, F. Herrera, and M. Lozano, \"Using evolutionary algorithms\nas instance selection for data reduction in kdd: an experimental study,\"\nIEEE Transactions on Evolutionary Computation, vol. 7, no. 6, pp.\n561\u2013575, 2004.\n[25] D. Borthakur, \"The hadoop distributed file system: Architecture and\ndesign,\" Hadoop Project Website, vol. 11, no. 11, pp. 1 \u2013 10, 2007.\n[26] D. B. Skalak, \"Prototype and feature selection by sampling and random\nmutation hill climbing algorithms,\" Machine Learning Proceedings, pp.\n293\u2013301, 1994.\n[27] V. S. Devi and M. N. Murty, \"An incremental prototype set building\ntechnique,\" Pattern Recognition, vol. 35, no. 2, pp. 505\u2013513, 2002."]}

Instance selection (IS) technique is used to reduce the data size to improve the performance of data mining methods. Recently, to process very large data set, several proposed methods divide the training set into some disjoint subsets and apply IS algorithms independently to each subset. In this paper, we analyze the limitation of these methods and give our viewpoint about how to divide and conquer in IS procedure. Then, based on fast condensed nearest neighbor (FCNN) rule, we propose a large data sets instance selection method with MapReduce framework. Besides ensuring the prediction accuracy and reduction rate, it has two desirable properties: First, it reduces the work load in the aggregation node; Second and most important, it produces the same result with the sequential version, which other parallel methods cannot achieve. We evaluate the performance of FCNN-MR on one small data set and two large data sets. The experimental results show that it is effective and practical.

Related Organizations
Keywords

Instance selection, data reduction, MapReduce, kNN.

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
    OpenAIRE UsageCounts
    Usage byUsageCounts
    visibility views 3
    download downloads 5
  • 3
    views
    5
    downloads
    Powered byOpenAIRE UsageCounts
Powered by OpenAIRE graph
Found an issue? Give us feedback
visibility
download
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
views
OpenAIRE UsageCountsViews provided by UsageCounts
downloads
OpenAIRE UsageCountsDownloads provided by UsageCounts
0
Average
Average
Average
3
5
Green