
Changelog The following listed datasets were recreated via a series of post-processing pipelines (available here) to eliminate a bias between the positive and negative classes (miRNA frequency class bias) discovered in previous versions of the datasets. All have a 1:1 positive to negative class ratio. AGO2_eCLIP_Manakov2022_leftout.tsv.gz AGO2_eCLIP_Manakov2022_test.tsv.gz AGO2_eCLIP_Manakov2022_train.tsv.gz AGO2_eCLIP_Klimentova2022.tsv.gz AGO2_CLASH_Hejret2023_test.tsv.gz AGO2_CLASH_Hejret2023_train.tsv.gz The structure of each dataset is consistent, with the following column order: gene: A string of length 50 indicating the binding site sequence in the 5’ to 3’ direction. noncodingRNA: A string of variable length (16–28) indicating the mature miRNA sequence in the 5’ to 3’ direction. noncodingRNA_name: A string indicating the name of the miRNA. noncodingRNA_fam: A string indicating the name of the miRNA family the miRNA belongs to. feature: A string indicating the feature annotation on the genome where the binding site occurs. label: A boolean value indicating whether the example belongs to the positive or negative class. chr: A string indicating the chromosome number on the genome where the binding site occurs. start: An integer indicating the 1-based start position of the binding site on the genome. end: An integer indicating the 1-based end position of the binding site on the genome. strand: A string indicating whether the binding site occurs on the ’+’ or ’-’ strand on the genome. gene_cluster_ID: An integer indicating the cluster ID of the binding site sequence used to generate the negative class. The following listed dataset is the concatenated HybriDetector output of all the selected samples from the available Manakov sample files. It therefore contains only a raw version of the positive class of the Manakov dataset. It is the input to the series of post-process pipelines for the Manakov dataset. AGO2_eCLIP_Manakov2022_full_dataset.tsv.gz The other inputs to the post-process pipelines for the Hejret and Klimentova datasets are found at the following links. Hejret dataset Klimentova dataset Note that the binding sites reported in all datasets are consistent with GRCh38.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
