Functional Inference from Orthology and Domain Architecture
Total Page:16
File Type:pdf, Size:1020Kb
Load more
Recommended publications
-
Applied Category Theory for Genomics – an Initiative
Applied Category Theory for Genomics { An Initiative Yanying Wu1,2 1Centre for Neural Circuits and Behaviour, University of Oxford, UK 2Department of Physiology, Anatomy and Genetics, University of Oxford, UK 06 Sept, 2020 Abstract The ultimate secret of all lives on earth is hidden in their genomes { a totality of DNA sequences. We currently know the whole genome sequence of many organisms, while our understanding of the genome architecture on a systematic level remains rudimentary. Applied category theory opens a promising way to integrate the humongous amount of heterogeneous informations in genomics, to advance our knowledge regarding genome organization, and to provide us with a deep and holistic view of our own genomes. In this work we explain why applied category theory carries such a hope, and we move on to show how it could actually do so, albeit in baby steps. The manuscript intends to be readable to both mathematicians and biologists, therefore no prior knowledge is required from either side. arXiv:2009.02822v1 [q-bio.GN] 6 Sep 2020 1 Introduction DNA, the genetic material of all living beings on this planet, holds the secret of life. The complete set of DNA sequences in an organism constitutes its genome { the blueprint and instruction manual of that organism, be it a human or fly [1]. Therefore, genomics, which studies the contents and meaning of genomes, has been standing in the central stage of scientific research since its birth. The twentieth century witnessed three milestones of genomics research [1]. It began with the discovery of Mendel's laws of inheritance [2], sparked a climax in the middle with the reveal of DNA double helix structure [3], and ended with the accomplishment of a first draft of complete human genome sequences [4]. -
Gene Prediction: the End of the Beginning Comment Colin Semple
View metadata, citation and similar papers at core.ac.uk brought to you by CORE provided by PubMed Central http://genomebiology.com/2000/1/2/reports/4012.1 Meeting report Gene prediction: the end of the beginning comment Colin Semple Address: Department of Medical Sciences, Molecular Medicine Centre, Western General Hospital, Crewe Road, Edinburgh EH4 2XU, UK. E-mail: [email protected] Published: 28 July 2000 reviews Genome Biology 2000, 1(2):reports4012.1–4012.3 The electronic version of this article is the complete one and can be found online at http://genomebiology.com/2000/1/2/reports/4012 © GenomeBiology.com (Print ISSN 1465-6906; Online ISSN 1465-6914) Reducing genomes to genes reports A report from the conference entitled Genome Based Gene All ab initio gene prediction programs have to balance sensi- Structure Determination, Hinxton, UK, 1-2 June, 2000, tivity against accuracy. It is often only possible to detect all organised by the European Bioinformatics Institute (EBI). the real exons present in a sequence at the expense of detect- ing many false ones. Alternatively, one may accept only pre- dictions scoring above a more stringent threshold but lose The draft sequence of the human genome will become avail- those real exons that have lower scores. The trick is to try and able later this year. For some time now it has been accepted increase accuracy without any large loss of sensitivity; this deposited research that this will mark a beginning rather than an end. A vast can be done by comparing the prediction with additional, amount of work will remain to be done, from detailing independent evidence. -
Meeting Review: Bioinformatics and Medicine – from Molecules To
Comparative and Functional Genomics Comp Funct Genom 2002; 3: 270–276. Published online 9 May 2002 in Wiley InterScience (www.interscience.wiley.com). DOI: 10.1002/cfg.178 Feature Meeting Review: Bioinformatics And Medicine – From molecules to humans, virtual and real Hinxton Hall Conference Centre, Genome Campus, Hinxton, Cambridge, UK – April 5th–7th Roslin Russell* MRC UK HGMP Resource Centre, Genome Campus, Hinxton, Cambridge CB10 1SB, UK *Correspondence to: Abstract MRC UK HGMP Resource Centre, Genome Campus, The Industrialization Workshop Series aims to promote and discuss integration, automa- Hinxton, Cambridge CB10 1SB, tion, simulation, quality, availability and standards in the high-throughput life sciences. UK. The main issues addressed being the transformation of bioinformatics and bioinformatics- based drug design into a robust discipline in industry, the government, research institutes and academia. The latest workshop emphasized the influence of the post-genomic era on medicine and healthcare with reference to advanced biological systems modeling and simulation, protein structure research, protein-protein interactions, metabolism and physiology. Speakers included Michael Ashburner, Kenneth Buetow, Francois Cambien, Cyrus Chothia, Jean Garnier, Francois Iris, Matthias Mann, Maya Natarajan, Peter Murray-Rust, Richard Mushlin, Barry Robson, David Rubin, Kosta Steliou, John Todd, Janet Thornton, Pim van der Eijk, Michael Vieth and Richard Ward. Copyright # 2002 John Wiley & Sons, Ltd. Received: 22 April 2002 Keywords: bioinformatics; -
Prospects & Overviews Orthology Prediction Methods: a Quality Assessment Using Curated Protein Families
Prospects & Overviews Methods, Models & Techniques Orthology prediction methods: A quality assessment using curated protein families Kalliopi Trachana1), Tomas A. Larsson1)2), Sean Powell1), Wei-Hua Chen1), Tobias Doerks1), Jean Muller3)4) and Peer Bork1)5)Ã The increasing number of sequenced genomes has Introduction prompted the development of several automated orthology prediction methods. Tests to evaluate the accuracy of pre- The analysis of fully sequenced genomes offers valuable insights into the function and evolution of biological systems dictions and to explore biases caused by biological and [1]. The annotation of newly sequenced genomes, comparative technical factors are therefore required. We used 70 man- and functional genomics, and phylogenomics depend on ually curated families to analyze the performance of five reliable descriptions of the evolutionary relationships of public methods in Metazoa. We analyzed the strengths protein families. All the members within a protein family and weaknesses of the methods and quantified the impact are homologous and can be further separated into orthologs, of biological and technical challenges. From the latter part which are genes derived through speciation from a single ancestral sequence, and paralogs, which are genes resulting of the analysis, genome annotation emerged as the largest from duplication events before and after speciation (out- and single influencer, affecting up to 30% of the performance. in-paralogy, respectively) [2, 3]. The large number of fully Generally, most methods did well in assigning orthologous sequenced genomes and the fundamental role of orthology group but they failed to assign the exact number of genes in modern biology have led to the development of a plethora of forhalfofthegroups.Thepublicly available benchmark set methods (e.g. -
UNIVERSITY of CALIFORNIA SAN DIEGO Making Sense of Microbial
UNIVERSITY OF CALIFORNIA SAN DIEGO Making sense of microbial populations from representative samples A dissertation submitted in partial satisfaction of the requirements for the degree Doctor of Philosophy in Computer Science by James T. Morton Committee in charge: Professor Rob Knight, Chair Professor Pieter Dorrestein Professor Rachel Dutton Professor Yoav Freund Professor Siavash Mirarab 2018 Copyright James T. Morton, 2018 All rights reserved. The dissertation of James T. Morton is approved, and it is acceptable in quality and form for publication on microfilm and electronically: Chair University of California San Diego 2018 iii DEDICATION To my friends and family who paved the road and lit the journey. iv EPIGRAPH The ‘paradox’ is only a conflict between reality and your feeling of what reality ‘ought to be’ —Richard Feynman v TABLE OF CONTENTS Signature Page . iii Dedication . iv Epigraph . .v Table of Contents . vi List of Abbreviations . ix List of Figures . .x List of Tables . xi Acknowledgements . xii Vita ............................................. xiv Abstract of the Dissertation . xvii Chapter 1 Methods for phylogenetic analysis of microbiome data . .1 1.1 Introduction . .2 1.2 Phylogenetic Inference . .4 1.3 Phylogenetic Comparative Methods . .6 1.4 Ancestral State Reconstruction . .9 1.5 Analysis of phylogenetic variables . 11 1.6 Using Phylogeny-Aware Distances . 15 1.7 Challenges of phylogenetic analysis . 18 1.8 Discussion . 19 1.9 Acknowledgements . 21 Chapter 2 Uncovering the horseshoe effect in microbial analyses . 23 2.1 Introduction . 24 2.2 Materials and Methods . 34 2.3 Acknowledgements . 35 Chapter 3 Balance trees reveal microbial niche differentiation . 36 3.1 Introduction . -
Methodology for Predicting Semantic Annotations of Protein Sequences by Feature Extraction Derived of Statistical Contact Potentials and Continuous Wavelet Transform
Universidad Nacional de Colombia Sede Manizales Master’s Thesis Methodology for predicting semantic annotations of protein sequences by feature extraction derived of statistical contact potentials and continuous wavelet transform Author: Supervisor: Gustavo Alonso Arango Dr. Cesar German Argoty Castellanos Dominguez A thesis submitted in fulfillment of the requirements for the degree of Master’s on Engineering - Industrial Automation in the Department of Electronic, Electric Engineering and Computation Signal Processing and Recognition Group June 2014 Universidad Nacional de Colombia Sede Manizales Tesis de Maestr´ıa Metodolog´ıapara predecir la anotaci´on sem´antica de prote´ınaspor medio de extracci´on de caracter´ısticas derivadas de potenciales de contacto y transformada wavelet continua Autor: Tutor: Gustavo Alonso Arango Dr. Cesar German Argoty Castellanos Dominguez Tesis presentada en cumplimiento a los requerimientos necesarios para obtener el grado de Maestr´ıaen Ingenier´ıaen Automatizaci´onIndustrial en el Departamento de Ingenier´ıaEl´ectrica,Electr´onicay Computaci´on Grupo de Procesamiento Digital de Senales Enero 2014 UNIVERSIDAD NACIONAL DE COLOMBIA Abstract Faculty of Engineering and Architecture Department of Electronic, Electric Engineering and Computation Master’s on Engineering - Industrial Automation Methodology for predicting semantic annotations of protein sequences by feature extraction derived of statistical contact potentials and continuous wavelet transform by Gustavo Alonso Arango Argoty In this thesis, a method to predict semantic annotations of the proteins from its primary structure is proposed. The main contribution of this thesis lies in the implementation of a novel protein feature representation, which makes use of the pairwise statistical contact potentials describing the protein interactions and geometry at the atomic level. -
Standardised Benchmarking in the Quest for Orthologs
View metadata, citation and similar papers at core.ac.uk brought to you by CORE provided by Harvard University - DASH Standardised Benchmarking in the Quest for Orthologs The Harvard community has made this article openly available. Please share how this access benefits you. Your story matters Citation Altenhoff, A. M., B. Boeckmann, S. Capella-Gutierrez, D. A. Dalquen, T. DeLuca, K. Forslund, J. Huerta-Cepas, et al. 2016. “Standardised Benchmarking in the Quest for Orthologs.” Nature methods 13 (5): 425-430. doi:10.1038/nmeth.3830. http://dx.doi.org/10.1038/ nmeth.3830. Published Version doi:10.1038/nmeth.3830 Citable link http://nrs.harvard.edu/urn-3:HUL.InstRepos:29408292 Terms of Use This article was downloaded from Harvard University’s DASH repository, and is made available under the terms and conditions applicable to Other Posted Material, as set forth at http:// nrs.harvard.edu/urn-3:HUL.InstRepos:dash.current.terms-of- use#LAA HHS Public Access Author manuscript Author ManuscriptAuthor Manuscript Author Nat Methods Manuscript Author . Author manuscript; Manuscript Author available in PMC 2016 October 04. Published in final edited form as: Nat Methods. 2016 May ; 13(5): 425–430. doi:10.1038/nmeth.3830. Standardised Benchmarking in the Quest for Orthologs Adrian M. Altenhoff1,2, Brigitte Boeckmann3, Salvador Capella-Gutierrez4,5,6, Daniel A. Dalquen7, Todd DeLuca8, Kristoffer Forslund9, Jaime Huerta-Cepas9, Benjamin Linard10, Cécile Pereira11,12, Leszek P. Pryszcz4, Fabian Schreiber13, Alan Sousa da Silva13, Damian Szklarczyk14,15, Clément-Marie Train1, Peer Bork9,16,17, Odile Lecompte18, Christian von Mering14,15, Ioannis Xenarios3,19,20, Kimmen Sjölander21, Lars Juhl Jensen22, Maria J. -
Complementary Symbiont Contributions to Plant Decomposition in a Fungus-Farming Termite
Complementary symbiont contributions to plant decomposition in a fungus-farming termite Michael Poulsena,1,2, Haofu Hub,1, Cai Lib,c, Zhensheng Chenb, Luohao Xub, Saria Otania, Sanne Nygaarda, Tania Nobred,3, Sylvia Klaubaufe, Philipp M. Schindlerf, Frank Hauserg, Hailin Panb, Zhikai Yangb, Anton S. M. Sonnenbergh, Z. Wilhelm de Beeri, Yong Zhangb, Michael J. Wingfieldi, Cornelis J. P. Grimmelikhuijzeng, Ronald P. de Vriese, Judith Korbf,4, Duur K. Aanend, Jun Wangb,j, Jacobus J. Boomsmaa, and Guojie Zhanga,b,2 aCentre for Social Evolution, Department of Biology, University of Copenhagen, DK-2100 Copenhagen, Denmark; bChina National Genebank, BGI-Shenzen, Shenzhen 518083, China; cCentre for GeoGenetics, Natural History Museum of Denmark, University of Copenhagen, DK-1350 Copenhagen, Denmark; dLaboratory of Genetics, Wageningen University, 6708 PB, Wageningen, The Netherlands; eFungal Biodiversity Centre, Centraalbureau voor Schimmelcultures, Royal Netherlands Academy of Arts and Sciences, NL-3584 CT, Utrecht, The Netherlands; fBehavioral Biology, Fachbereich Biology/Chemistry, University of Osnabrück, D-49076 Osnabrück, Germany; gCenter for Functional and Comparative Insect Genomics, Department of Biology, University of Copenhagen, DK-2100 Copenhagen, Denmark; hDepartment of Plant Breeding, Wageningen University and Research Centre, NL-6708 PB, Wageningen, The Netherlands; iDepartment of Microbiology, Forestry and Agricultural Biotechnology Institute, University of Pretoria, Pretoria SA-0083, South Africa; and jDepartment of Biology, University of Copenhagen, DK-2100 Copenhagen, Denmark Edited by Ian T. Baldwin, Max Planck Institute for Chemical Ecology, Jena, Germany, and approved August 15, 2014 (received for review October 24, 2013) Termites normally rely on gut symbionts to decompose organic levels-of-selection conflicts that need to be regulated (12). -
Peregrine and Saker Falcon Genome Sequences Provide Insights Into Evolution of a Predatory Lifestyle
LETTERS OPEN Peregrine and saker falcon genome sequences provide insights into evolution of a predatory lifestyle Xiangjiang Zhan1,7, Shengkai Pan2,7, Junyi Wang2,7, Andrew Dixon3, Jing He2, Margit G Muller4, Peixiang Ni2, Li Hu2, Yuan Liu2, Haolong Hou2, Yuanping Chen2, Jinquan Xia2, Qiong Luo2, Pengwei Xu2, Ying Chen2, Shengguang Liao2, Changchang Cao2, Shukun Gao2, Zhaobao Wang2, Zhen Yue2, Guoqing Li2, Ye Yin2, Nick C Fox3, Jun Wang5,6 & Michael W Bruford1 As top predators, falcons possess unique morphological, 16,263 genes were predicted for F. peregrines, and 16,204 were pre- physiological and behavioral adaptations that allow them to be dicted for F. cherrug (Supplementary Table 6 and Supplementary successful hunters: for example, the peregrine is renowned as Note). Approximately 92% of these genes were functionally annotated the world’s fastest animal. To examine the evolutionary basis of using homology-based methods (Supplementary Table 7). predatory adaptations, we sequenced the genomes of both the Comparative genome analysis was carried out to assess evolution peregrine (Falco peregrinus) and saker falcon (Falco cherrug), and innovation within falcons using related genomes with comparable All rights reserved. and we present parallel, genome-wide evidence for evolutionary assembly quality (Supplementary Table 8). Orthologous genes were innovation and selection for a predatory lifestyle. The genomes, identified in the chicken, zebra finch, turkey, peregrine and saker assembled using Illumina deep sequencing with greater than using the program TreeFam3. For the five genome-enabled avian spe- 100-fold coverage, are both approximately 1.2 Gb in length, cies, a maximum-likelihood phylogeny using 861,014 4-fold degen- with transcriptome-assisted prediction of approximately 16,200 erate sites from 6,267 single-copy orthologs confirmed that chicken America, Inc. -
Multiscale Modeling in Systems Biology
Digital Comprehensive Summaries of Uppsala Dissertations from the Faculty of Science and Technology 2051 Multiscale Modeling in Systems Biology Methods and Perspectives ADRIEN COULIER ACTA UNIVERSITATIS UPSALIENSIS ISSN 1651-6214 ISBN 978-91-513-1225-5 UPPSALA URN urn:nbn:se:uu:diva-442412 2021 Dissertation presented at Uppsala University to be publicly examined in 2446 ITC, Lägerhyddsvägen 2, Uppsala, Friday, 10 September 2021 at 10:15 for the degree of Doctor of Philosophy. The examination will be conducted in English. Faculty examiner: Professor Mark Chaplain (University of St Andrews). Abstract Coulier, A. 2021. Multiscale Modeling in Systems Biology. Methods and Perspectives. Digital Comprehensive Summaries of Uppsala Dissertations from the Faculty of Science and Technology 2051. 60 pp. Uppsala: Acta Universitatis Upsaliensis. ISBN 978-91-513-1225-5. In the last decades, mathematical and computational models have become ubiquitous to the field of systems biology. Specifically, the multiscale nature of biological processes makes the design and simulation of such models challenging. In this thesis we offer a perspective on available methods to study and simulate such models and how they can be combined to handle biological processes evolving at different scales. The contribution of this thesis is threefold. First, we introduce Orchestral, a multiscale modular framework to simulate multicellular models. By decoupling intracellular chemical kinetics, cell-cell signaling, and cellular mechanics by means of operator-splitting, it is able to combine existing software into one massively parallel simulation. Its modular structure makes it easy to replace its components, e.g. to adjust the level of modeling details. We demonstrate the scalability of our framework on both high performance clusters and in a cloud environment. -
EMBO Encounters Issue43.Pdf
WINTER 2019/2020 ISSUE 43 Nine group leaders selected Meet the first EMBO Global Investigators PAGE 6 Accelerating scientific publishing EMBO publishing costs Review Commons Making our journals’ platform announced finances public PAGE 3 PAGES 10 – 11 Welcome, Young Investigators! Contract replaces stipend Marking ten years 27 group leaders join the programme EMBO Postdoctoral Fellowships EMBO Molecular Medicine receive an update celebrates anniversary PAGES 4 – 5 PAGE 7 PAGE 13 www.embo.org TABLE OF CONTENTS EMBO NEWS EMBO news Review Commons: accelerating publishing Page 3 EMBO Molecular Medicine turns ten © Marietta Schupp, EMBL Photolab Marietta Schupp, © Page 13 Editorial MBO was founded by scientists for Introducing 27 new Young Investigators scientists. This philosophy remains at Pages 4-5 Ethe heart of our organization until today. EMBO Members are vital in the running of our Meet the first EMBO Global programmes and activities: they screen appli- Accelerating scientific publishing 17 journals on board Investigators cations, interview candidates, decide on fund- Review Commons will manage the transfer of ing, and provide strategic direction. On pages EMBO and ASAPbio announced pre-journal portable review platform the manuscript, reviews, and responses to affili- Page 6 8-9 four members describe why they chose to ate journals. A consortium of seventeen journals New members meet in Heidelberg dedicate their time to an EMBO Committee across six publishers (see box) have joined the Fellowships: from stipends to contracts Pages 14 – 15 and what they took away from the experience. n December 2019, EMBO, in partnership with decide to submit their work to a journal, it will project by committing to use the Review Commons Page 7 When EMBO was created, the focus lay ASAPbio, launched Review Commons, a multi- allow editors to make efficient editorial decisions referee reports for their independent editorial deci- specifically on fostering cross-border inter- Ipublisher partnership which aims to stream- based on existing referee comments. -
Assigning Folds to the Proteins Encoded by the Genome of Mycoplasma Genitalium (Protein Fold Recognition͞computer Analysis of Genome Sequences)
Proc. Natl. Acad. Sci. USA Vol. 94, pp. 11929–11934, October 1997 Biophysics Assigning folds to the proteins encoded by the genome of Mycoplasma genitalium (protein fold recognitionycomputer analysis of genome sequences) DANIEL FISCHER* AND DAVID EISENBERG University of California, Los Angeles–Department of Energy Laboratory of Structural Biology and Molecular Medicine, Molecular Biology Institute, University of California, Los Angeles, Box 951570, Los Angeles, CA 90095-1570 Contributed by David Eisenberg, August 8, 1997 ABSTRACT A crucial step in exploiting the information genitalium (MG) (10), as a test of the capabilities of our inherent in genome sequences is to assign to each protein automatic fold recognition server and as a case study to sequence its three-dimensional fold and biological function. identify the difficulties facing automated fold assignment. Here we describe fold assignment for the proteins encoded by the small genome of Mycoplasma genitalium. The assignment MATERIALS AND METHODS was carried out by our computer server (http:yywww.doe- mbi.ucla.eduypeopleyfrsvryfrsvr.html), which assigns folds to The MG Sequences. The 468 MG sequences were obtained amino acid sequences by comparing sequence-derived predic- from The Institute for Genome Research (TIGR) through its tions with known structures. Of the total of 468 protein ORFs, Web address: http:yywww.tigr.orgytdbymdbymgdbymgd- 103 (22%) can be assigned a known protein fold with high b.html. Three types of annotation (based on searches in the confidence, as cross-validated with tests on known structures. sequence database) accompany each TIGR sequence (10): (i) Of these sequences, 75 (16%) show enough sequence similarity functional assignment—a clear sequence similarity with a to proteins of known structure that they can also be detected protein of known function from another organism was found by traditional sequence–sequence comparison methods.