
The DTI dataset introduced in the CoNCISE paper. https://www.biorxiv.org/content/10.1101/2025.01.08.632039 We additionally add an extended dataset that includes additional DTI information acquired from PLINDER. moodeng-v1.tar.gz entries: train.csv (number of positives approximately 110000) test.csv val.csv These files are comma separated CSV files. Each file has four fields: Target Sequence: The protein sequence input Label: A binary value. 1 indicates that the sequence binds with the ligand. O otherwise SMILES: The ligand smiles Remapped Entry: The representative sequence of the cluster the `Target Sequence` belongs to. Discard this field while doing DTI prediction moodeng-v2-extended.tar.gz entries: This is an updated dataset that with around 4 million entries (including the negative examples, negative:positive = 10:1; number of positives in train approximately 450000). The archive file contains two TAB separated CSV files: train.csv and test.csv Following are their fields: SMILES: same as above Target Sequence: same as above Label: Same as above ID: A short name used to represent the `Target sequence` cluster-assignment: The integer ID representing the cluster assignment.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
