MACHOS: Markov clusters of homologous subsequences

descriptionPublicationkeyboard_double_arrow_right Article , Conference object 01 Jul 2008 Australia English Publisher:Oxford University Press (OUP)Journal:Bioinformatics, volume 24, pages i77-i85 (issn: 1367-4803, eissn: 1367-4811,

Copyright policy )Funded by:ARC | Australian Centre for Gen...

Authors: Simon Wong; Mark A. Ragan;

doi: 10.1093/bioinformatics/btn144

pmid: 18586748

pmc: PMC2718622

MACHOS: Markov clusters of homologous subsequences

- Summary
- Subjects
- Metrics

Abstract

Abstract Motivation: The classification of proteins into homologous groups (families) allows their structure and function to be analysed and compared in an evolutionary context. The modular nature of eukaryotic proteins presents a considerable challenge to the delineation of families, as different local regions within a single protein may share common ancestry with distinct, even mutually exclusive, sets of homologs, thereby creating an intricate web of homologous relationships if full-length sequences are taken as the unit of evolution. We attempt to disentangle this web by developing a fully automated pipeline to delineate protein subsequences that represent sensible units for homology inference, and clustering them into putatively homologous families using the Markov clustering algorithm. Results: Using six eukaryotic proteomes as input, we clustered 162 349 protein sequences into 19 697–77 415 subsequence families depending on granularity of clustering. We validated these Markov clusters of homologous subsequences (MACHOS) against the manually curated Pfam domain families, using a quality measure to assess overlap. Our subsequence families correspond well to known domain families and achieve higher quality scores than do groups generated by fully automated domain family classification methods. We illustrate our approach by analysis of a group of proteins that contains the glutamyl/glutaminyl-tRNA synthetase domain, and conclude that our method can produce high-coverage decomposition of protein sequence space into precise homologous families in a way that takes the modularity of eukaryotic proteins into account. This approach allows for a fine-scale examination of evolutionary histories of proteins encoded in eukaryotic genomes. Contact: m.ragan@imb.uq.edu.au Supplementary information: Supplementary data are available at Bioinformatics online. MACHOS for the six proteomes are available as FASTA-formatted files: http://research1t.imb.uq.edu.au/ragan/machos

Country

Australia

Related Organizations

Keywords

Biochemistry & Molecular Biology, Proteome, Statistics & Probability, Molecular Sequence Data, 069999 Biological Sciences not elsewhere classified, Biochemical Research Methods, C1, Ismb 2008 Conference Proceedings 19–23 July 2008, Toronto, Sequence Analysis, Protein, Cluster Analysis, Interdisciplinary Applications, Computer Simulation, Amino Acid Sequence, Models, Statistical, Sequence Homology, Amino Acid, 006, BIOTECHNOLOGY & APPLIED MICROBIOLOGY, Markov Chains, Protein Structure, Tertiary, MATHEMATICAL & COMPUTATIONAL BIOLOGY, Biotechnology & Applied Microbiology, Models, Chemical, 970106 Expanding Knowledge in the Biological Sciences, Computer Science, Computer Science, Interdisciplinary Applications, Mathematical & Computational Biology, Sequence Alignment, Mathematics, BIOCHEMICAL RESEARCH METHODS, Algorithms

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	12
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%