
Abstract The era of biodiversity genomics is characterized by large-scale genome sequencing efforts that aim to represent each living taxon with an assembled genome. Generating knowledge from this wealth of data has not kept up with this pace. We here discuss major challenges to integrating these novel genomes into a comprehensive functional and evolutionary network spanning the tree of life. In summary, the expanding datasets create a need for scalable gene annotation methods. To trace gene function across species, new methods must seek to increase the resolution of ortholog analyses, e.g. by extending analyses to the protein domain level and by accounting for alternative splicing. Additionally, the scope of orthology prediction should be pushed beyond well-investigated proteomes. This demands the development of specialized methods for the identification of orthologs to short proteins and noncoding RNAs and for the functional characterization of novel gene families. Furthermore, protein structures predicted by machine learning are now readily available, but this new information is yet to be integrated with orthology-based analyses. Finally, an increasing focus should be placed on making orthology assignments adhere to the findable, accessible, interoperable, and reusable (FAIR) principles. This fosters green bioinformatics by avoiding redundant computations and helps integrating diverse scientific communities sharing the need for comparative genetics and genomics information. It should also help with communicating orthology-related concepts in a format that is accessible to the public, to counteract existing misinformation about evolution.
Genomics/methods; Biodiversity; Animals; Evolution, Molecular; Molecular Sequence Annotation; Computational Biology/methods; FAIR; annotation transfer; domain architecture; noncoding RNA; ortholog search; protein structure, 570, domain architecture, Computational Biology, Molecular Sequence Annotation, Review, Genomics, Biodiversity, Annotation transfer, noncoding RNA, Noncoding RNA, Evolution, Molecular, Ortholog search, annotation transfer; domain architecture; FAIR; noncoding RNA; ortholog search; protein structure, Domain architecture, Protein structure, Animals, annotation transfer, protein structure, ortholog search, FAIR
Genomics/methods; Biodiversity; Animals; Evolution, Molecular; Molecular Sequence Annotation; Computational Biology/methods; FAIR; annotation transfer; domain architecture; noncoding RNA; ortholog search; protein structure, 570, domain architecture, Computational Biology, Molecular Sequence Annotation, Review, Genomics, Biodiversity, Annotation transfer, noncoding RNA, Noncoding RNA, Evolution, Molecular, Ortholog search, annotation transfer; domain architecture; FAIR; noncoding RNA; ortholog search; protein structure, Domain architecture, Protein structure, Animals, annotation transfer, protein structure, ortholog search, FAIR
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 23 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
