Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Dataset . 2025
License: CC BY
Data sources: ZENODO
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Dataset . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Dataset . 2025
License: CC BY
Data sources: Datacite
ZENODO
Dataset . 2026
License: CC BY
Data sources: Datacite
ZENODO
Dataset . 2025
License: CC BY
Data sources: Datacite
versions View all 3 versions
addClaim

PanTEon Database: A Cross-Kingdom, Automatically Curated Reference for Transposable Elements

Authors: Orozco Arias, Simon; Ferrer-Pomer, Iamil; Rodrigues de Góes, Fabiana; Gaviria-Orrego, Simon; Rossi Paschoal, Alexandre; guyot, romain; Gabaldón, Toni; +2 Authors

PanTEon Database: A Cross-Kingdom, Automatically Curated Reference for Transposable Elements

Abstract

The PanTEon Database is a freely available collection of more than 126,000 automatically curated transposable element (TE) sequences, spanning animals, plants, and fungi and covering all major TE orders. The database was designed to maximize sequence fidelity, taxonomic diversity, and methodological consistency, making it suitable for both training and benchmarking state-of-the-art TE classification tools. Data sources and integration PanTEon integrates TE sequences from multiple complementary resources: Curated sequences from Dfam (version 3.9) Automatically curated sequences from APTEdb (Pedro et al., 2021) Uncurated sequences from Dfam TE sequences from the Ensembl 2023 release (Martin et al., 2023) All uncurated sequences were originally generated using RepeatModeler2 (Flynn et al., 2020) and subsequently automatically curated with MCHelper (Orozco-Arias et al., 2024). Only TEs showing clear evidence of structural completeness and expected length profiles were retained, resulting in a high-confidence dataset suitable for training and benchmarking machine learning models. Sequence identification Each TE sequence in the PanTEon Database is assigned a standardized identifier composed of: A sequence name (either provided by Dfam or systematically assigned by the PanTEon framework), A three-level classification (Class / Order / Superfamily), The species of origin. For example, a TE sequence derived from curated Dfam data is identified as: >PumCon-1.141#CLASSII/TIR/TC1MARINER @Puma concolor In this case, PumCon-1.141 is the original family name, the element belongs to the TIR/Tc1–Mariner superfamily, and it was obtained from Puma concolor. In contrast, a TE sequence that was originally uncurated and subsequently processed and integrated into the PanTEon Database follows the systematic naming scheme: >PDB00000038#CLASSI/LTR/LARD @Certhia brachydactyla Metadata and taxonomic context Additional taxonomic information—such as order, family, phylum, and higher ranks—as well as details about the origin of each sequence, are provided in the accompanying metadata file: PanTEon_Database_metadata_v.1.5.1.csv Benchmark edition To enable fair and reproducible benchmarking of state-of-the-art TE classification tools, a benchmark edition of the PanTEon Database was generated. This version includes only TE superfamilies represented by more than 10 sequences and merges rare or uncommon superfamilies with their closest relatives (see the PanTEon paper for full details). The benchmark dataset is available as: PanTEon_Database_v1.5.1_benchmark_edition.fasta Trained models for PanTEon Inference This repository also contains pre-trained models corresponding to the different in-built architectures of PanTEon (classification task). These models can be downloaded and used with the PanTEon inference module by specifying their paths via the -d parameter. This release includes models trained on the following datasets: all (entire PanTEon Database v1.5.1), Animalia, Chordata, Arthropods, Plantae, Angiosperms, and Fungi. PanTEon can be downloaded from the following GitHub repository: https://github.com/simonorozcoarias/PanTEon Why PanTEon? By combining broad taxonomic coverage, automated curation, and a standardized nomenclature, the PanTEon Database provides a robust reference resource for: developing and training machine learning and deep learning models, benchmarking TE classification tools, and exploring TE diversity across kingdoms. Funding Simon Orozco-Arias is supported by a fellowship within the “Generación D” initiative, Red.es, Ministerio para la Transformación Digital y de la Función Pública, for talent attraction (C005/24-ED CV1). Funded by the European Union NextGenerationEU funds, through PRTR.

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average