Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Dataset . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Dataset . 2026
License: CC BY
Data sources: Datacite
ZENODO
Dataset . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

An Imbalanced Dataset with Multiple Feature Sets for Studying Quality Control of Next-Generation Sequencing

Authors: Röchner, Philipp; Krämer, Clarissa; Mayer, Johannes U; Rothlauf, Franz; Albrecht, Steffen; Sprang, Maximilian;

An Imbalanced Dataset with Multiple Feature Sets for Studying Quality Control of Next-Generation Sequencing

Abstract

Next-generation sequencing (NGS) is a key technique for studying the DNA and RNA of organisms. However, identifying diverse quality problems in NGS dataacross different experimental settings remains challenging. To develop automated quality-control tools, researchers require datasets with features that capture thecharacteristics of quality problems. Existing NGS repositories, however, offer only a limited number of quality-related features. To address this gap, we proposea dataset derived from 37,491 NGS samples with two types of quality-related feature representations. The first type consists of 34 features derived from qualitycontrol tools (QC-34 features). The second type has a variable number of features ranging from eight to 1,183. These features were derived from read counts inproblematic genomic regions identified by the ENCODE blocklist (BL features) [1]. All feature sets are tabular and describe the quality of the same 37,491 human and mouse samples, but capture different aspects of sample quality. Based on automated quality control and a manual review by domain experts, 3.2% of the samples are classified as low quality and labeled as revoked; the remaining samples are high quality and labeled as released. You can find our paper here: https://arxiv.org/abs/2604.04981 QC-34 Features The QC-34 features consist of the raw (RAW), mapping (MAP), transcription start site (TSS), and location (LOC) features, as introduced by Albrecht et al. [2] . In total, there are 34 features. The RAW features are ordinal; all other QC-34 features are numeric. The QC-34.csv file contains the 34 quality-related features for the 37,491 NGS samples and their quality labels. The first part of the QC-34 feature names refers to the corresponding feature set (RAW, MAP, TSS, and LOC). The second part describes the quality metric. BL Features The BL-n.csv files contain the quality-related features for the 37,491 NGS samples and their quality labels, where the number n refers to the number of features.The names of the BL features encode three types of information. The first two letters indicate whether the blocklisted region is from a human (hs) or a mouse (mm). The next two to three capital letters indicate whether the region is low mappability (LM) or a high-signal region (HSR) according to the ENCODE blocklist. The final number, separated by an underscore, refers to the genomic region in the original ENCODE blocklist. For example, the hsHSR_17 feature describes the number of reads mapped to the 17th blocklisted region of the ENCODE blocklist for humans. Sample Metadata The fastq_samples_meta.csv file contains metadata features of the FASTQ samples derived from ENCODE. For example, the Accession feature of each of the 37,491 samples in the QC-34.csv, and BL-n.csv contains an ID referring to the corresponding ENCODE sample. Experiment Metadata The experiments_meta.csv file contains the metadata of the ENCODE experiments from which we took the FASTQ samples. For example, the Lab features contain the laboratory that provided the data. Donor Metadata The donor_ethnicity.csv, donor_sex.csv, and donor_life_stage.csv files provide information about the donors from whom the samples in our datasets were obtained.

Keywords

machine learning, Next-generation sequencing, data quality, imbalanced data, bioinformatics, benchmarking, quality control, outlier detection, anomaly detection

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average