bioRxiv preprint doi: https://doi.org/10.1101/2020.04.30.068296; this version posted August 6, 2020. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under aCC-BY 4.0 International license. OMAmer: tree-driven and alignment-free protein assignment to subfamilies outperforms closest sequence approaches 4,5,* Victor Rossier1, 2,3, Alex Warwick Vesztrocy1, 2, 3, Marc Robinson-Rechavi and Christophe Dessimoz1, 2, 3, 5,6,* 1Department of Computational Biology, University of Lausanne, Switzerland; 2Center for Integrative Genomics, University of Lausanne, Switzerland; 3SIB Swiss Institute of Bioinformatics, Lausanne, Switzerland; 4Department of Ecology and Evolution, University of Lausanne, Switzerland; 5Department of Genetics, Evolution, and Environment, University College London, UK; 6Department of Computer Science, University College London, UK. *Corresponding authors:
[email protected] &
[email protected] Abstract Assigning new sequences to known protein Families and subFamilies is a prerequisite For many functional, comparative and evolutionary genomics analyses. Such assignment is commonly achieved by looking For the closest sequence in a reFerence database, using a method such as BLAST. However, ignoring the gene phylogeny can be misleading because a query sequence does not necessarily belong to the same subFamily as its closest sequence. For example, a hemoglobin which branched out prior to the hemoglobin alpha/beta duplication could be closest to a hemoglobin alpha or beta sequence, whereas it is neither. To overcome this problem, phylogeny-driven tools have emerged but rely on gene trees, whose inFerence is computationally expensive.