
doi: 10.2139/ssrn.6906059
Resolving deep phylogenetic relationships within Coleoptera remains challenging despite advances in phylogenomics. We present an updated phylogenomic dataset integrating available assemblies, reprocessed Sequence Read Archive (SRA) data, and newly generated genomes. The dataset comprises 263 beetle species (104 families and 210 subfamilies) and 13 outgroups, based on 2,280 orthologous loci selected from a Coleoptera-specific BUSCO set. A multi-step pipeline optimised orthology assignment, alignment quality, and locus selection, incorporating sequence- and tree-based filtering criteria to reduce bias from compositional heterogeneity, rate variation, and model violations. Phylogenetic analyses of concatenated datasets under multiple substitution models and recoding schemes showed high congruence for moderately complete matrices (50–70% taxon occupancy), whereas stringent matrices (≥90%) reduced accuracy due to data loss. The preferred topology obtained under the PMSF model supports the basal split between Polyphaga and other suborders, and recovers key early-diverging polyphagan lineages. Relationships among infraorders are largely consistent with previous genomic studies, although uncertainty persists within Elateroidea, Staphylinidae, Curculionidae, and the superfamily-level relationships within Cucujiformia, while our framework extends to the subfamily level, unlike previous genomic studies. Our stringent locus filtering improved phylogenetic signal and concordance, although the excluded loci recover broadly similar topologies. The GHOST mixture model yielded a slightly different, plausible topology, while coalescent approaches in ASTRAL were less satisfactory. Gene concordance factors were used to identify the remaining weakly supported nodes across the tree. Our study underscores the importance of balanced taxon sampling, rigorous data curation, and model selection, and establishes the foundation for future studies of Coleoptera diversification.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
