bioRxiv preprint doi: https://doi.org/10.1101/115774; this version posted June 4, 2018. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under aCC-BY-NC 4.0 International license. Running head: LIKELIHOOD AND OUTLIERS IN PHYLOGENOMICS Title: Analyzing contentious relationships and outlier genes in phylogenomics Joseph F. Walker1*, Joseph W. Brown2, and Stephen A. Smith1* 1Deptartment Ecology and Evolutionary Biology, University of Michigan, Ann Arbor, Michigan, 48109, USA 2Department of Animal and Plant Sciences, University of Sheffield, Sheffield, S10 2TN, United Kingdom *Corresponding authors Corresponding author emails:
[email protected],
[email protected] 1 bioRxiv preprint doi: https://doi.org/10.1101/115774; this version posted June 4, 2018. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under aCC-BY-NC 4.0 International license. ABSTRACT Recent studies have demonstrated that conflict is common among gene trees in phylogenomic studies, and that less than one percent of genes may ultimately drive species tree inference in supermatrix analyses. Here, we examined two datasets where supermatrix and coalescent-based species trees conflict. We identified two highly influential “outlier” genes in each dataset. When removed from each dataset, the inferred supermatrix trees matched the topologies obtained from coalescent analyses. We also demonstrate that, while the outlier genes in the vertebrate dataset have been shown in a previous study to be the result of errors in orthology detection, the outlier genes from a plant dataset did not exhibit any obvious systematic error and therefore may be the result of some biological process yet to be determined.