
Distinguishing genuine microbial proteins from contaminants remains a major bottleneck in genomics, particularly for environmental and non-model organisms where conventional homology-based tools are slow, resource-intensive, and leave large fractions of the "dark proteome" unclassified. LA4SR offers a scalable, interpretable framework that classifies algal and bacterial proteins directly from translated sequence data, achieving near-complete recall while accelerating inference by ~ 10,000-fold relative to BLASTP. By revealing that internal sequence features alone can drive robust classification, LA4SR bypasses the need for complete gene models or perfect annotations—opening new opportunities for analyzing complex microbial communities and metagenomes. Interpretability methods further link emergent amino acid signatures to evolutionary and ecological features, highlighting the potential of language models not only to accelerate genomics workflows but also to uncover new biological insights. Table S1 | External spreadsheet. This spreadsheet contains LA4SR performance metrics, technical performance estimations, and BLAST results and runtimes of genomes comprising the algal training data. Table S2 | External spreadsheet. Captum attributions for 100 sequences each of algal and bacterial origin obtained using the LayerIntegratedGradients function. Table S3 | External spreadsheet. Influential motifs found with the DeepMotifMinerPro software introduced in this work (see Data S3). Table S4 | External spreadsheet. LA4SR and Diamond BLAST results for data from new assemblies from seen species (Fig. S7), contaminated assemblies from unseen genera (Fig. S7), and clean assemblies from unseen genera (Fig. 8). For the LA4SR results for genome assemblies from unseen genera, the genomes were published after the model was trained, and the genera shown were not included in the training dataset. Newly sequenced genomes uploaded to NCBI SRA accession SUB14799921.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
