Protein classification with imbalanced data

descriptionPublicationkeyboard_double_arrow_right Article 12 Dec 2007 English Publisher:WileyJournal:Proteins: Structure, Function, and Bioinformatics, volume 70, pages 1,125-1,132 (issn: 0887-3585, eissn: 1097-0134,

Copyright policy )

Authors: Xing-Ming, Zhao; Xin, Li; Luonan, Chen; Kazuyuki, Aihara;

doi: 10.1002/prot.21870

pmid: 18076026

Protein classification with imbalanced data

- Summary
- Subjects
- Metrics

Abstract

AbstractGenerally, protein classification is a multi‐class classification problem and can be reduced to a set of binary classification problems, where one classifier is designed for each class. The proteins in one class are seen as positive examples while those outside the class are seen as negative examples. However, the imbalanced problem will arise in this case because the number of proteins in one class is usually much smaller than that of the proteins outside the class. As a result, the imbalanced data cause classifiers to tend to overfit and to perform poorly in particular on the minority class.This article presents a new technique for protein classification with imbalanced data. First, we propose a new algorithm to overcome the imbalanced problem in protein classification with a new sampling technique and a committee of classifiers. Then, classifiers trained in different feature spaces are combined together to further improve the accuracy of protein classification. The numerical experiments on benchmark datasets show promising results, which confirms the effectiveness of the proposed method in terms of accuracy. The Matlab code and supplementary materials are available at http://eserver2.sat.iis.u‐tokyo.ac.jp/∼xmzhao/proteins.html. Proteins 2008. © 2007 Wiley‐Liss, Inc.

Related Organizations

Shanghai University
China (People's Republic of)
Hong Kong Baptist University
China (People's Republic of)
Osaka Sangyo University
Japan
Chinese Academy of Sciences
China (People's Republic of)
University of Tokyo
Japan

View all View all

Keywords

Internet, Artificial Intelligence, Discriminant Analysis, Proteins, Classification, Databases, Protein, Algorithms

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	111
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 1%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 1%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%