MEGAN analysis of metagenomic data

descriptionPublicationkeyboard_double_arrow_right Article 25 Jan 2007 English Publisher:Cold Spring Harbor LaboratoryJournal:Genome Research, volume 17, pages 377-386 (issn: 1088-9051,

Copyright policy )

Authors: Daniel H, Huson; Alexander F, Auch; Ji, Qi; Stephan C, Schuster;

doi: 10.1101/gr.5969107

pmid: 17255551

pmc: PMC1800929

MEGAN analysis of metagenomic data

- Summary
- Subjects
- Metrics

Abstract

Metagenomics is the study of the genomic content of a sample of organisms obtained from a common habitat using targeted or random sequencing. Goals include understanding the extent and role of microbial diversity. The taxonomical content of such a sample is usually estimated by comparison against sequence databases of known sequences. Most published studies use the analysis of paired-end reads, complete sequences of environmental fosmid and BAC clones, or environmental assemblies. Emerging sequencing-by-synthesis technologies with very high throughput are paving the way to low-cost random “shotgun” approaches. This paper introduces MEGAN, a new computer program that allows laptop analysis of large metagenomic data sets. In a preprocessing step, the set of DNA sequences is compared against databases of known sequences using BLAST or another comparison tool. MEGAN is then used to compute and explore the taxonomical content of the data set, employing the NCBI taxonomy to summarize and order the results. A simple lowest common ancestor algorithm assigns reads to taxa such that the taxonomical level of the assigned taxon reflects the level of conservation of the sequence. The software allows large data sets to be dissected without the need for assembly or the targeting of specific phylogenetic markers. It provides graphical and statistical output for comparing different data sets. The approach is applied to several data sets, including the Sargasso Sea data set, a recently published metagenomic data set sampled from a mammoth bone, and several complete microbial genomes. Also, simulations that evaluate the performance of the approach for different read lengths are presented.

Related Organizations

University of Tübingen
Germany
University of Tuebingen
Germany
Pennsylvania State University
United States

Keywords

Genome, Species Specificity, Computational Biology, Genetic Variation, Biodiversity, Genomics, Atlantic Ocean, Ecosystem, Phylogeny, Software

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	3K
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 0.01%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 0.01%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 0.1%

Found an issue? Give us feedback

3K

Top 0.01%

Top 0.1%

bronze

Fields of Science (3) View all

medical and health sciences

basic medicine

Fields of Science

medical and health sciences

basic medicine

View all