Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions

Name: Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions
Keywords: Papyrus, machine learning, bioactivity, cheminformatics, proteochemometrics

Béquignon, Olivier J. M.; Bongers, Brandon J.; Jespers, Willem; IJzerman, Ad P.; van de Water, Bob; van Westen, Gerard J. P.

Found an issue? Give us feedback

ZENODOarrow_drop_down

ZENODO

Dataset . 2022

License: CC BY SA

Data sources: Datacite

ZENODO

Dataset . 2022

License: CC BY SA

Data sources: Datacite

ZENODO

Dataset . 2024

License: CC BY SA

Data sources: Datacite

ZENODO

Dataset . 2024

License: CC BY SA

Data sources: Datacite

Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions

Research datakeyboard_double_arrow_right Dataset 24 Aug 2022Publisher:ZenodoFunded by:EC | eTRANSAFE

Authors: Béquignon, Olivier J. M.; Bongers, Brandon J.; Jespers, Willem; IJzerman, Ad P.; van de Water, Bob; van Westen, Gerard J. P.;

doi: 10.5281/zenodo.7821773 , 10.5281/zenodo.10943207 , 10.5281/zenodo.7019873 , 10.5281/zenodo.7019874

Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions

- Summary
- Subjects
- Related research
  (9)
- Metrics

Abstract

Fixed version of Papyrus++ 05.5: - In the previous 05.5 version data was incorrectly duplicated based on assay type. This resulted in unintended data augmentation. - In this fixed 05.5 version the duplicates have been eliminated, now reporting the correct amount of data per assay type. This repository contains the version 05.5 of the Papyrus dataset, an aggregated dataset of small molecule bioactivities, as described in the article "Papyrus - A large scale curated dataset aimed at bioactivity predictions" http://doi.org/10.1186/s13321-022-00672-x. With the ongoing rapid growth of publicly available ligand-protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers’ time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million datapoints. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some example quantitative structure-activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing a solid baseline for related research.

Related Organizations

Leiden University
Netherlands

Keywords

Papyrus, machine learning, bioactivity, cheminformatics, proteochemometrics

9 Research products, page 1 of 1

Papyrus: a large-scale curated dataset aimed at bioactivity predictions
2021IsNewVersionOf
DrugEx pretrained model (SMILES-based; RECAP; Papyrus 05.5)
2023IsSourceOf
DrugEx RNN-GRU pretrained model (Papyrus 05.5)
2023IsSourceOf
Dataset - Papyrus 05.4 - A large scale curated dataset aimed at bioactivity predictions
2022IsNewVersionOf
Dataset - Papyrus - A large scale curated dataset aimed at bioactivity predictions
2022IsPreviousVersionOf
DrugEx v2 pretrained model (Papyrus 05.5)
2022IsSourceOf
DrugEx v3 pretrained model (graph-based; Papyrus 05.5)
2022IsSourceOf
DrugEx pretrained model (graph-based; RECAP; Papyrus 05.5)
2023IsSourceOf
DrugEx pretrained model (SMILES-based; Papyrus 05.5)
2023IsSourceOf

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	2
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average