
pmid: 15059830
Abstract Motivation: Current DNA sequencing technology produces reads of about 500–750 bp, with typical coverage under 10×. New sequencing technologies are emerging that produce shorter reads (length 80–200 bp) but allow one to generate significantly higher coverage (30× and higher) at low cost. Modern assembly programs and error correction routines have been tuned to work well with current read technology but were not designed for assembly of short reads. Results: We analyze the limitations of assembling reads generated by these new technologies and present a routine for base-calling in reads prior to their assembly. We demonstrate that while it is feasible to assemble such short reads, the resulting contigs will require significant (if not prohibitive) finishing efforts. Availability: Available from the web at http://www.cse.ucsd.edu/groups/bioinformatics/software.html
Contig Mapping, Base Sequence, Gene Expression Profiling, Molecular Sequence Data, Feasibility Studies, Sequence Analysis, DNA, Sequence Alignment, Algorithms
Contig Mapping, Base Sequence, Gene Expression Profiling, Molecular Sequence Data, Feasibility Studies, Sequence Analysis, DNA, Sequence Alignment, Algorithms
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 147 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 1% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 0.1% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
