
Next-generation sequencing (NGS) is a key technique for studying the DNA and RNA of organisms. However, identifying diverse quality problems in NGS dataacross different experimental settings remains challenging. To develop automated quality-control tools, researchers require datasets with features that capture thecharacteristics of quality problems. Existing NGS repositories, however, offer only a limited number of quality-related features. To address this gap, we proposea dataset derived from 37,491 NGS samples with two types of quality-related feature representations. The first type consists of 34 features derived from qualitycontrol tools (QC-34 features). The second type has a variable number of features ranging from eight to 1,183. These features were derived from read counts inproblematic genomic regions identified by the ENCODE blocklist (BL features) [1]. All feature sets are tabular and describe the quality of the same 37,491 human and mouse samples, but capture different aspects of sample quality. Based on automated quality control and a manual review by domain experts, 3.2% of the samples are classified as low quality and labeled as revoked; the remaining samples are high quality and labeled as released. You can find our paper here: https://arxiv.org/abs/2604.04981 QC-34 Features The QC-34 features consist of the raw (RAW), mapping (MAP), transcription start site (TSS), and location (LOC) features, as introduced by Albrecht et al. [2] . In total, there are 34 features. The RAW features are ordinal; all other QC-34 features are numeric. The QC-34.csv file contains the 34 quality-related features for the 37,491 NGS samples and their quality labels. The first part of the QC-34 feature names refers to the corresponding feature set (RAW, MAP, TSS, and LOC). The second part describes the quality metric. BL Features The BL-n.csv files contain the quality-related features for the 37,491 NGS samples and their quality labels, where the number n refers to the number of features.The names of the BL features encode three types of information. The first two letters indicate whether the blocklisted region is from a human (hs) or a mouse (mm). The next two to three capital letters indicate whether the region is low mappability (LM) or a high-signal region (HSR) according to the ENCODE blocklist. The final number, separated by an underscore, refers to the genomic region in the original ENCODE blocklist. For example, the hsHSR_17 feature describes the number of reads mapped to the 17th blocklisted region of the ENCODE blocklist for humans. Sample Metadata The fastq_samples_meta.csv file contains metadata features of the FASTQ samples derived from ENCODE. For example, the Accession feature of each of the 37,491 samples in the QC-34.csv, and BL-n.csv contains an ID referring to the corresponding ENCODE sample. Experiment Metadata The experiments_meta.csv file contains the metadata of the ENCODE experiments from which we took the FASTQ samples. For example, the Lab features contain the laboratory that provided the data. Donor Metadata The donor_ethnicity.csv, donor_sex.csv, and donor_life_stage.csv files provide information about the donors from whom the samples in our datasets were obtained.
machine learning, Next-generation sequencing, data quality, imbalanced data, bioinformatics, benchmarking, quality control, outlier detection, anomaly detection
machine learning, Next-generation sequencing, data quality, imbalanced data, bioinformatics, benchmarking, quality control, outlier detection, anomaly detection
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
