Papyrus: a large-scale curated dataset aimed at bioactivity predictions

Name: Papyrus: a large-scale curated dataset aimed at bioactivity predictions
Keywords: FOS: Computer and information sciences, Papyrus, curated dataset, Medicinal and Biomolecular Chemistry, Artificial Intelligence and Image Processing, cheminformatics tool, Theoretical and Computational Chemistry, FOS: Chemical sciences, bioactivity data, machine Learning Predictions

BÃ©quignon, Olivier; Bongers, Brandon; Jespers, W. (Willem); IJzerman, Adriaan P.; van de Water, Bob; Van westen, Gerard JP

Found an issue? Give us feedback

figsharearrow_drop_down

figshare

Collection . 2024

License: CC BY

Data sources: Datacite

figshare

Collection . 2024

License: CC BY

Data sources: Datacite

4TU.ResearchData

Dataset . 2021

License: CC BY

Data sources: 4TU.ResearchData

4TU.ResearchData

Dataset . 2022

License: CC BY SA

Data sources: 4TU.ResearchData

4TU.ResearchData

Dataset . 2022

License: CC BY SA

Data sources: 4TU.ResearchData

4TU.ResearchData

Dataset . 2021

License: CC BY

Data sources: 4TU.ResearchData

4TU.ResearchData | science.engineering.design

Dataset . 2021

License: CC BY

Data sources: Datacite

4TU.ResearchData | science.engineering.design

Dataset . 2021

License: CC BY

Data sources: Datacite

4TU.ResearchData | science.engineering.design

Dataset . 2022

License: CC BY SA

Data sources: Datacite

4TU.ResearchData | science.engineering.design

Dataset . 2022

License: CC BY SA

Data sources: Datacite

Papyrus: a large-scale curated dataset aimed at bioactivity predictions

Research datakeyboard_double_arrow_right Collection , Dataset 29 Oct 2021Publisher:figshare

Authors: BÃ©quignon, Olivier; Bongers, Brandon; Jespers, W. (Willem); IJzerman, Adriaan P.; van de Water, Bob; Van westen, Gerard JP;

doi: 10.6084/m9.figshare.c.6579035.v1 , 10.4121/16896406.v2 , 10.4121/16896406.v1 , 10.4121/16896406 , 10.6084/m9.figshare.c.6579035 , 10.4121/16896406.v3

Papyrus: a large-scale curated dataset aimed at bioactivity predictions

- Summary
- Subjects
- Related research
  (3)
- Metrics

Abstract

Abstract With the ongoing rapid growth of publicly available ligand–protein bioactivity data, there is a trove of valuable data that can be used to train a plethora of machine-learning algorithms. However, not all data is equal in terms of size and quality and a significant portion of researchers’ time is needed to adapt the data to their needs. On top of that, finding the right data for a research question can often be a challenge on its own. To meet these challenges, we have constructed the Papyrus dataset. Papyrus is comprised of around 60 million data points. This dataset contains multiple large publicly available datasets such as ChEMBL and ExCAPE-DB combined with several smaller datasets containing high-quality data. The aggregated data has been standardised and normalised in a manner that is suitable for machine learning. We show how data can be filtered in a variety of ways and also perform some examples of quantitative structure–activity relationship analyses and proteochemometric modelling. Our ambition is that this pruned data collection constitutes a benchmark set that can be used for constructing predictive models, while also providing an accessible data source for research. Graphical Abstract

Related Organizations

Leiden University
Netherlands

Keywords

FOS: Computer and information sciences, Papyrus, curated dataset, Medicinal and Biomolecular Chemistry, Artificial Intelligence and Image Processing, cheminformatics tool, Theoretical and Computational Chemistry, FOS: Chemical sciences, bioactivity data, machine Learning Predictions

3 Research products, page 1 of 1

Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions
2022IsPreviousVersionOf
Dataset - Papyrus 05.4 - A large scale curated dataset aimed at bioactivity predictions
2022IsSupplementedBy
Dataset - Papyrus - A large scale curated dataset aimed at bioactivity predictions
2022IsPreviousVersionOf

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	1
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

1

Average

Related to Research communities

Netherlands Research Portal

Papyrus: a large-scale curated dataset aimed at bioactivity predictions

Papyrus: a large-scale curated dataset aimed at bioactivity predictions

3 Research products, page 1 of 1

Accompanying data - Papyrus - A large scale curated dataset aimed at bioactivity predictions

Dataset - Papyrus 05.4 - A large scale curated dataset aimed at bioactivity predictions

Dataset - Papyrus - A large scale curated dataset aimed at bioactivity predictions