On the complexity of haplotyping a microbial community

descriptionPublicationkeyboard_double_arrow_right Article , Other literature type 10 Aug 2020 Belgium, United Kingdom English Publisher:Oxford University Press (OUP)Journal:Bioinformatics, volume 37, pages 1,360-1,366 (issn: 1367-4803, eissn: 1367-4811,

Copyright policy )Funded by:EC | MASTER

Authors: Samuel M. Nicholls; Wayne Aubrey; Kurt De Grave; Leander Schietgat; Christopher J. Creevey; Amanda Clare;

doi: 10.1093/bioinformatics/btaa977 , 10.1101/2020.08.10.244848

pmid: 33444437

pmc: PMC8208737

On the complexity of haplotyping a microbial community

- Summary
- Subjects
- Metrics

Abstract

Abstract Motivation Population-level genetic variation enables competitiveness and niche specialization in microbial communities. Despite the difficulty in culturing many microbes from an environment, we can still study these communities by isolating and sequencing DNA directly from an environment (metagenomics). Recovering the genomic sequences of all isoforms of a given gene across all organisms in a metagenomic sample would aid evolutionary and ecological insights into microbial ecosystems with potential benefits for medicine and biotechnology. A significant obstacle to this goal arises from the lack of a computationally tractable solution that can recover these sequences from sequenced read fragments. This poses a problem analogous to reconstructing the two sequences that make up the genome of a diploid organism (i.e. haplotypes) but for an unknown number of individuals and haplotypes. Results The problem of single individual haplotyping was first formalized by Lancia et al. in 2001. Now, nearly two decades later, we discuss the complexity of ‘haplotyping’ metagenomic samples, with a new formalization of Lancia et al.’s data structure that allows us to effectively extend the single individual haplotype problem to microbial communities. This work describes and formalizes the problem of recovering genes (and other genomic subsequences) from all individuals within a complex community sample, which we term the metagenomic individual haplotyping problem. We also provide software implementations for a pairwise single nucleotide variant (SNV) co-occurrence matrix and greedy graph traversal algorithm. Availability and implementation Our reference implementation of the described pairwise SNV matrix (Hansel) and greedy haplotype path traversal algorithm (Gretel) is open source, MIT licensed and freely available online at github.com/samstudio8/hansel and github.com/samstudio8/gretel, respectively.

Countries

Belgium, United Kingdom

Related Organizations

KU Leuven
Belgium
Vrije Universiteit Brussel
Vrije Universiteit Brussel
Belgium
Queen's University
Canada
AGRIFOOD AND BIOSCIENCES INSTITUTE
United Kingdom

View all View all

Keywords

Statistics and Probability, 570, Technology, Biochemistry & Molecular Biology, Bioinformatics, Statistics & Probability, GENOMES, Biochemistry, 46 Information and computing sciences, Biochemical Research Methods, Molecular Biology, 01 Mathematical Sciences, Science & Technology, 31 Biological sciences, 06 Biological Sciences, Original Papers, 004, Computer Science Applications, Computational Mathematics, Computational Theory and Mathematics, Biotechnology & Applied Microbiology, Physical Sciences, Computer Science, Computer Science, Interdisciplinary Applications, Mathematical & Computational Biology, 08 Information and Computing Sciences, Life Sciences & Biomedicine, 49 Mathematical sciences, Mathematics

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	19
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%