
Supplementary Table 1. High-quality genomes with species assignment and classification in the Representative and Clonal Group datasets. Supplementary Table 2. Metadata and genomic parameter values for all high-quality Representative genomes. Supplementary Table 3. Phenotypic and ecological data integrated from three publicly available data sources. Supplementary Table 4. Species-level genomic parameter ranks and values for 15,235 species (NCBI taxonomy). Each table entry shows the rank (rank/total species) and the median±SD. Supplementary Table 5. Species-level population genetic parameter ranks and values for 786 species (NCBI taxonomy). Each table entry shows the rank (rank/total species) and the median±SD. Supplementary Table 6. Correlation matrices among genomic, population genetic, phylogenetic, phenotypic, and ecological parameters based on the NCBI taxonomy. a) Correlation matrices for all individual parameters. b) Correlation matrices between parameter categories, where each category–category correlation is represented by the maximum correlation coefficient among all parameter pairs within the two categories. Supplementary Table 7. Species-level genomic parameter ranks and values for 42,469 species (GTDB taxonomy). Each table entry shows the rank (rank/total species) and the median±SD. Supplementary Table 8. Species-level population genetic parameter ranks and values for 1411 species (GTDB taxonomy). Each table entry shows the rank (rank/total species) and the median±SD. Supplementary Table 9. Correlation matrices among genomic and population genetic parameters based on the GTDB taxonomy. MAP_rawdata.tar.gz Whole Genome Alignments This folder contains two subfolders: "Representative Dataset" and "Clonal Group Dataset". "Representative Dataset": Contains whole genome alignments for representative genomes of species with ≥10 representative strains. The files are named as “Species_taxid.whole_genome_aln.fasta”. "Clonal Group Dataset": Contains whole genome alignments for clonal groups with ≥10 strains. The files are named as “Species_taxid.CGid.whole_genome_aln.fasta”. Each genome is named by its "biosample-accession". Data in both datasets are compressed in 7z format. SNP Matrix This folder also contains two subfolders: "Representative Dataset" and "Clonal Group Dataset". These subfolders include core-genome SNP matrix files corresponding to the whole genome alignments in the datasets mentioned above. The files are named as “Species_taxid.snp.matrix” (Representative Dataset) or “Species_taxid.CGid.snp.matrix” (Clonal Group Dataset). In each SNP matrix, rows represent SNP positions in the reference genome, and columns represent the strains. Data are compressed in 7z format. Pangenome This folder contains the presence/absence matrix of pan-genes (generated by Panaroo analysis) for species with ≥10 representative strains. In each matrix, rows represent pan-genes, and columns represent the representative strains. Gene presence is indicated by “1”, and gene absence by “0”. The files are named as “Species_taxid.gene_presence_absence.Rtab”. Data are compressed in gzip format.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
