Efficient computation of statistics for words with mismatches

descriptionPublicationkeyboard_double_arrow_right Article , Conference object 01 Jan 2010Embargo end date: 23 Aug 2010 Germany, Italy English Publisher:Schloss Dagstuhl – Leibniz-Zentrum für Informatik

Authors: PIZZI, CINZIA;

doi: 10.4230/dagsemproc.10231.4

handle: 11577/2420299

Efficient computation of statistics for words with mismatches

- Summary
- Subjects
- Metrics

Abstract

Since early stages of bioinformatics, substrings played a crucial role in the search and discovery of significant biological signals. Despite the advent of a large number of different approaches and models toaccomplish these tasks, substrings continue to be widely used to determine statistical distributions and compositions of biological sequences at various levels of details. Here we overview efficient algorithms that were recently proposed to compute the actual and the expected frequency for words with k mismatches, when it is assumed that the words of interest occur at least once exactly in the sequence under analysis. Efficiency means these algorithms are polynomial in k rather than exponential as with an enumerative approach, and independent on the length of the query word. These algorithms are all based on a common incremental approach of a preprocessing step that allows to answer queries related to any word occurring in the text efficiently. The same approach can be used with a sliding window scanning of the sequence to compute the same statistics for words of fixed lengths, even more efficiently. The efficient computation of both expected and actual frequency of sub- strings, combined with a study on the monotonicity of popular scores such as z-scores, allows to build tables of feasible size in reasonable time, and can therefore be used in practical applications.

Countries

Germany, Italy

Related Organizations

University of Padua
Italy
Leibniz Association
Germany
Schloss Dagstuhl – Leibniz Center for Informatics
Germany
Helsinki Institute for Information Technology
Finland

Keywords

dynamic programming, biological sequences, mismatches, biological sequences., Statistics on words, 004

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

0

Average

Green

Fields of Science (4) View all

engineering and technology

medical engineering

Fields of Science

engineering and technology

medical engineering

View all