descriptionPublicationkeyboard_double_arrow_right Article 09 Dec 2006 English Publisher:American Chemical Society (ACS)Journal:Journal of Chemical Information and Modeling, volume 47, pages 25-33 (issn: 1549-9596, eissn: 1549-960X,

Authors: James L, Melville; Jenna F, Riley; Jonathan D, Hirst;

doi: 10.1021/ci600384z

pmid: 17238245

Similarity by Compression

- Summary
- Subjects
- Metrics

Abstract

We present a simple and effective method for similarity searching in virtual high-throughput screening, requiring only a string-based representation of the molecules (e.g., SMILES) and standard compression software, available on all modern desktop computers. This method utilizes the normalized compression distance, an approximation of the normalized information distance, based on the concept of Kolmogorov complexity. On representative data sets, we demonstrate that compression-based similarity searching can outperform standard similarity searching protocols, exemplified by the Tanimoto coefficient combined with a binary fingerprint representation and data fusion. Software to carry out compression-based similarity is available from our Web site at http://comp.chem.nottingham.ac.uk/download/zippity.

Related Organizations

University of Nottingham
United Kingdom

Keywords

Databases, Factual, Molecular Structure, Area Under Curve, Organic Chemicals, Data Compression, Software

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	13
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%

Found an issue? Give us feedback

Average

Top 10%

Fields of Science (4) View all

Fields of Science

Upload OA version

Are you the author of this publication? Upload your Open Access version to Zenodo!

It’s fast and easy, just two clicks!

uploadUpload now