aaHash: recursive amino acid sequence hashing

descriptionPublicationkeyboard_double_arrow_right Article , Other literature type 01 Jan 2023Publisher:Cold Spring Harbor LaboratoryJournal:Bioinformatics Advances, volume 3 (eissn: 2635-0041,

Copyright policy )Funded by:CIHR | unidentified, NIH | De Novo Assembly Tools: R...

Authors: Johnathan Wong; Parham Kazemi; Lauren Coombe; René L Warren; Inanç Birol;

doi: 10.1101/2023.05.08.539909 , 10.1093/bioadv/vbad162

pmid: 37214907 , 38023332

pmc: PMC10197579 , PMC10660294

aaHash: recursive amino acid sequence hashing

- Summary
- Subjects
- Related research
  (2)
- Metrics

Abstract

AbstractMotivationK-mer hashing is a common operation in many foundational bioinformatics problems. However, generic string hashing algorithms are not optimized for this application. Strings in bioinformatics use specific alphabets, a trait leveraged for nucleic acid sequences in earlier work. We note that amino acid sequences, with complexities and context that cannot be captured by generic hashing algorithms, can also benefit from a domain-specific hashing algorithm. Such a hashing algorithm can accelerate and improve the sensitivity of bioinformatics applications developed for protein sequences.ResultsHere, we present aaHash, a recursive hashing algorithm tailored for amino acid sequences. This algorithm utilizes multiple hash levels to represent biochemical similarities between amino acids. aaHash performs ∼10X faster than generic string hashing algorithms in hashing adjacentk-mers.Availability and implementationaaHash is available online athttps://github.com/bcgsc/btlliband is free for academic use.

Related Organizations

Canada's Michael Smith Genome Sciences Centre
Canada

Keywords

Application Note, Article

2 Research products, page 1 of 1

aahash_paper software on GitHub
IsRelatedTo
btllib software on GitHub
IsRelatedTo

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

0

Average

Green

gold

Funded by

CIHR| unidentified, NIH| De Novo Assembly Tools: Research with Unbiased Engines - Renewal (DNA-TRUER)

aaHash: recursive amino acid sequence hashing

aaHash: recursive amino acid sequence hashing

2 Research products, page 1 of 1

aahash_paper software on GitHub

btllib software on GitHub