Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml Jakob Voss, based on art designer at PLoS, modified by Wikipedia users Nina and Beao Closed Access logo, derived from PLoS Open Access logo. This version with transparent background. http://commons.wikimedia.org/wiki/File:Closed_Access_logo_transparent.svg Jakob Voss, based on art designer at PLoS, modified by Wikipedia users Nina and Beao ZENODOarrow_drop_down
image/svg+xml Jakob Voss, based on art designer at PLoS, modified by Wikipedia users Nina and Beao Closed Access logo, derived from PLoS Open Access logo. This version with transparent background. http://commons.wikimedia.org/wiki/File:Closed_Access_logo_transparent.svg Jakob Voss, based on art designer at PLoS, modified by Wikipedia users Nina and Beao
ZENODO
Software . 2025
License: CC BY
Data sources: ZENODO
ZENODO
Software . 2025
License: CC BY
Data sources: Datacite
ZENODO
Software . 2025
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

mtDNA variant calling pipeline

Authors: Barra Matos, Gustavo; Araújo, Gilderlanio;

mtDNA variant calling pipeline

Abstract

FASTQ Processing and Analysis Pipeline Overview This Bash script automates the processing of paired-end FASTQ files, including quality assessment, trimming, alignment, and duplicate marking. It leverages tools like FastQC, MultiQC, Fastp, BWA, SAMtools, and Picard to perform the following steps: Quality Analysis using FastQC and MultiQC (Pre- and Post-trimming). Trimming and Adapter Removal using Fastp. Alignment of trimmed reads against a mitochondrial reference genome using BWA. Conversion of SAM to BAM, sorting, and duplicate marking using SAMtools and Picard. Prerequisites Ensure the following software/tools are installed and accessible in your $PATH: FastQC MultiQC Fastp BWA SAMtools Picard Java Runtime Environment (for Picard) Directory Structure Before running the script, create and organize your directories as follows: Input Directory: Contains raw FASTQ files (*_R1.fastq and *_R2.fastq). Output Directories: The script will automatically create these: Processed FASTQ files Quality reports (pre- and post-trimming) SAM/BAM files and final sorted BAM files Marked duplicate files and metrics How to Use Set Variables: Update the following paths in the script to match your file structure: input_dir: Path to raw FASTQ files. output_dir: Directory for trimmed FASTQ files. fastqc_dir_pre: Directory for pre-trimming quality reports. fastqc_dir_post: Directory for post-trimming quality reports. multiqc_dir: Directory for consolidated MultiQC reports. alignment_dir: Directory for alignment (SAM) files. reference_mtDNA: Path to the reference genome file (FASTA format). picard_output_dir: Directory for Picard-marked BAM files. Make Executable: chmod +x mitogenome_processing.sh Run the Script**: ./mitogenome_processing.sh Pipeline Workflow Initial Quality Check: Generates pre-trimming quality reports for raw FASTQ files using FastQC. Combines individual reports using MultiQC. Trimming with Fastp: Removes adapters and low-quality bases. Produces trimmed FASTQ files with a summary HTML report. Post-Quality Check: Generates quality reports for trimmed FASTQ files. Alignment: Aligns trimmed reads to the mitochondrial reference genome using BWA and outputs SAM files. Conversion and Sorting: Converts SAM files to BAM format using SAMtools. Sorts BAM files for downstream processing. Duplicate Marking: Marks and removes duplicates from BAM files using Picard. Outputs Quality Reports: Available in the directories for both pre- and post-trimming. Processed Files: Trimmed FASTQ files in output_dir. Aligned SAM files in alignment_dir. Sorted BAM files in sorted_bam. Marked BAM files and metrics in picard_output_dir. Troubleshooting Ensure paths to tools and reference files are correct. Check the output logs for errors during alignment or duplicate marking. Acknowledgments This pipeline uses tools developed by open-source communities. Please refer to their documentation for further details: FastQC MultiQC Fastp BWA SAMtools Picard References FastQC: Andrews, S. FastQC: A Quality Control Tool for High Throughput Sequence Data. Available at: FastQC MultiQC: Ewels, P., Magnusson, M., Lundin, S., & Käller, M. (2016). "MultiQC: summarize analysis results for multiple tools." Bioinformatics, 32(19), 3047–3048. https://doi.org/10.1093/bioinformatics/btw354 Fastp: Chen, S., Zhou, Y., Chen, Y., & Gu, J. (2018). "fastp: an ultra-fast all-in-one FASTQ processor." Bioinformatics, 34(17), i884–i890. https://doi.org/10.1093/bioinformatics/bty560 BWA: Li, H., & Durbin, R. (2009). "Fast and accurate short read alignment with Burrows-Wheeler transform." Bioinformatics, 25(14), 1754–1760. https://doi.org/10.1093/bioinformatics/btp324 SAMtools: Li, H., Handsaker, B., Wysoker, A., & Fennell, T. (2009). "The Sequence Alignment/Map format and SAMtools." Bioinformatics, 25(16), 2078–2079. https://doi.org/10.1093/bioinformatics/btp352 Picard: Broad Institute. (2020). Picard Tools. Available at: Picard

Related Organizations
  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average