Bioinformatics Tools for Mass Spectrometry, Phylogenetic Footprinting, and the Integration of Biological Data

Total Page:16

File Type:pdf, Size:1020Kb

Bioinformatics Tools for Mass Spectrometry, Phylogenetic Footprinting, and the Integration of Biological Data Bioinformatics tools for mass spectrometry, phylogenetic footprinting, and the integration of biological data Dissertation zur Erlangung des akademischen Grades Doktor der Naturwissenschaften (Dr.rer.nat.) der Naturwissenschaftliche Fakultät III Agrar- und Ernährungswissenschaften, Geowissenschaften und Informatik der Martin-Luther-Universität Halle-Wittenberg vorgelegt von Hendrik Treutler Geb. am 10.03.1985 in Karl–Marx–Stadt Gutachter: 1. Prof. Dr. Ivo Grosse 2. Prof. Dr. Burkhard Morgenstern 3. Dr. Steffen Neumann Tag der Verteidigung: 16. November 2017 ii Acknowledgements My first thanks belong to my beloved wife Anna-Maria Treutler. Thanks to you I can submit this thesis before our fifth wedding anniversary. You encourage me all the time and you gave me three wonderful children. Many thanks also belong to my parents who formed my roots in this world and made it possible for me to fulfill my wishes. I thank my instructors Ivo Grosse, Steffen Neumann, Stefan Posch, and Falk Schreiber for guiding me to this thesis. You formed my roots in the scientific world and your advice on technical, scientific, and political issues lead to own experiences which will be indispensable for me in the future. Martin Nettling earns special thanks. The exchange of thoughts with you concerning technical, scientific, and private issues changed and widened my mind countless times. With him on your side things will basically work out. I thank Gerd Balcke for accompanying me in the world of mass spectrometry. Your constructive ideas and your eagerness are an inspiration for me and lead to a great collaboration. I thank Karin Gorzolka for her passion to share thoughts in many positive discussions. I thank Susann Mönchgesang, Christoph Ruttkies, Daniel Schober, and Sarah Scharfenberg for their great support in technical and scientific issues and for their friendship. I thank Jesus Cerquides, Tobias Czauderna, Jan Grau, Eva Grafahrend-Belau, Anja Hartmann, Astrid Junker, Jens Keilwagen, Matthias Klapperstück, Matthias Lange, Christian Klukas, Hendrik Rohn, and Uwe Scholz for fruitful discussions and professional teamwork. Last but not least I thank Kristian Peters for technical support with MetFamily. You do a really great job and I enjoy your presence for many reasons. iv Contents 1 Summary 1 1.1 English version . .1 1.2 German version . .4 2 Introduction 7 2.1 Biological background . .9 2.1.1 The central dogma of molecular biology . .9 2.1.2 Regulation of gene expression and transcriptional initiation . 11 2.1.3 Metabolism and small molecules . 11 2.2 Computer science background . 13 2.2.1 The Java programming language . 14 2.2.2 The R programming language . 14 2.2.3 Databases . 15 2.2.4 Data integration and Ontologies . 16 2.3 Bioinformatics background . 16 2.3.1 Integration of biological data . 16 2.3.2 Phylogenetic footprinting and phylogenetic shadowing . 17 2.3.3 Mass spectrometry . 18 2.4 Research objectives . 19 3 Paper summary 23 3.1 The presented works in context . 23 3.2 Integration of biological data . 26 3.2.1 VANTED v2 – A framework for systems biology applications . 26 3.2.2 DBE2 – Management of experimental data for the VANTED system 28 3.3 Phylogenetic footprinting . 29 3.3.1 Detecting and correcting the binding–affinity bias in ChIP-seq data using inter–species information . 30 3.3.2 DiffLogo: a comparative visualization of sequence motifs . 31 3.4 Data Processing and Interpretation of Mass Spectrometry Data . 33 v CONTENTS 3.4.1 Discovering Regulated Metabolite Families in Untargeted Metabolomics Studies . 33 3.4.2 Prediction, Detection, and Validation of Isotope Clusters in Mass Spectrometry Data . 35 4 Integration of biological data 49 4.1 VANTED v2: a framework for systems biology applications . 49 4.2 DBE2 – Management of experimental data for the VANTED system . 63 5 Phylogenetic footprinting 75 5.1 Detecting and correcting the binding-affinity bias in ChIP-seq data using inter-species information . 75 5.2 DiffLogo: a comparative visualization of sequence motifs . 86 6 Data Processing and Interpretation of Mass Spectrometry Data 97 6.1 Discovering Regulated Metabolite Families in Untargeted Metabolomics Stud- ies ......................................... 97 6.2 Prediction, Detection, and Validation of Isotope Clusters in Mass Spectrom- etry Data . 107 vi 1 Summary 1.1 English version Complex biological processes such as RNA splicing, protein degradation, and metabolic switches are a prerequisite for organisms to adapt to environmental stimuli, respond to pathogens, and absorb nutrients. These biological processes are studied in molecular biol- ogy and there are powerful techniques for the study of proteins, DNA sequences, and small molecules such as nuclear magnetic resonance (NMR) spectroscopy, high–throughput se- quencing technologies, and mass spectrometry. Depending on the technique the amount of biological raw data can be in the order of hundreds of gigabytes and specialized tools are needed to process this data and extract interpretable information. A deep understanding of both biology and computer science is a prerequisite for the development of such tools and bioinformatics is an interdisciplinary field which emerged from this demand. Bioin- formatics research involves the development of tools for the study of biological processes using methods from computer science. Systems biology is a bioinformatics field which aims at an understanding of biological systems rather than investigating each biological process individually. Systems biology approaches are able to translate various kinds of biological data to parameters of mathe- matical models enabling, amongst others, the understanding of the complex interplay of different biological entities, the reproduction of emergent properties of metabolic pathways, and the prediction of the behaviour of cells under hypothetical conditions. Research in sys- tems biology necessitates insights into the regulation of gene expression, the metabolism, and the integration of biological data from different sources. The publications in this cumu- lative thesis are rooted in these three fields, namely (i) “Phylogenetic footprinting” for the study of the regulation of transcriptional initiation, (ii) “Data Processing and Interpretation of Mass Spectrometry Data” for the study of metabolic fingerprints, and (iii) the “Integra- tion of biological data” for the combined analysis of different data. Two publications reside in each of these fields and I will summarize these six publications subsequently. Decoding the regulation of gene expression is essential for the understanding of life and 1 1. SUMMARY the transcriptional initiation is a vital sub–process. Transcriptional initiation is governed by the concerted binding of transcription factors (TFs) to transcription factor binding sites (TFBSs) and de novo motif discovery with “Phylogenetic footprinting” on basis of ChIP–seq data gives deep insights into this sub–process. Unfortunately, ChIP–seq data is subject to the ubiquitous binding–affinity bias which leads to the prediction of distorted sequence motifs and biased sets of TFBSs. My colleagues and I developed a phylogenetic footprinting model which is capable of estimating and correcting the binding–affinity bias in ChIP–seq data. We found that the proposed phylogenetic footprinting model improves de novo motif discovery and that the corrected sequence motifs are softer than the uncorrected sequence motifs. The increasing availability of sequence data and algorithms for de novo motif discovery re- sults in the demand to compare different sequence motifs. Sequence motifs which have been extracted from, e.g., different species, tissues, and samples are often similar which makes the comparison challenging. My colleagues and I developed the R–based tool DiffLogo for the comparative visualization of sequence motifs in DNA sequences, RNA sequences, and protein domains. DiffLogo allows the detection of even small motif differences between two sequence motifs and also supports the pair–wise comparison of more than two sequence motifs. We demonstrated the utility of DiffLogo using sequence motifs of one transcription factor from different cell lines, sequence motifs from different transcription factors, and sequence motifs in protein domains from different phylogenetic kingdoms. DiffLogo was also used to show the effect of the binding–affinity bias in ChIP–seq data. The study of small molecules in the cell gives insights into the chemical fingerprints of different biological processes. Metabolite profiles provide a snapshot of these chemical fingerprints in the samples of interest and Mass Spectrometry (MS) is an essential technol- ogy to prepare these metabolite profiles. The “Data Processing and Interpretation of Mass Spectrometry Data” is mandatory in order to gain biological insights from MS data and the identification and precise quantification of metabolites remains a major challenge. Isotope clusters are sets of related signals in MS data and it has been shown that isotope clusters can be utilized to improve the identification and quantification of metabolites. My colleagues and I developed an approach for the prediction, detection, and validation of isotope clusters in MS data and we integrated this approach into the popular R–based tools xcms and CAMERA. We found that using the proposed approach it is possible to extract 37% more isotope signals from Arabidopsis thaliana measurements and in a mix of standard compounds the correct molecular formula could be predicted in 92% of the cases in the top three ranks. The structural elucidation
Recommended publications
  • A Method to Infer Changed Activity of Metabolic Function from Transcript Profiles
    ModeScore: A Method to Infer Changed Activity of Metabolic Function from Transcript Profiles Andreas Hoppe and Hermann-Georg Holzhütter Charité University Medicine Berlin, Institute for Biochemistry, Computational Systems Biochemistry Group [email protected] Abstract Genome-wide transcript profiles are often the only available quantitative data for a particular perturbation of a cellular system and their interpretation with respect to the metabolism is a major challenge in systems biology, especially beyond on/off distinction of genes. We present a method that predicts activity changes of metabolic functions by scoring reference flux distributions based on relative transcript profiles, providing a ranked list of most regulated functions. Then, for each metabolic function, the involved genes are ranked upon how much they represent a specific regulation pattern. Compared with the naïve pathway-based approach, the reference modes can be chosen freely, and they represent full metabolic functions, thus, directly provide testable hypotheses for the metabolic study. In conclusion, the novel method provides promising functions for subsequent experimental elucidation together with outstanding associated genes, solely based on transcript profiles. 1998 ACM Subject Classification J.3 Life and Medical Sciences Keywords and phrases Metabolic network, expression profile, metabolic function Digital Object Identifier 10.4230/OASIcs.GCB.2012.1 1 Background The comprehensive study of the cell’s metabolism would include measuring metabolite concentrations, reaction fluxes, and enzyme activities on a large scale. Measuring fluxes is the most difficult part in this, for a recent assessment of techniques, see [31]. Although mass spectrometry allows to assess metabolite concentrations in a more comprehensive way, the larger the set of potential metabolites, the more difficult [8].
    [Show full text]
  • The Kyoto Encyclopedia of Genes and Genomes (KEGG)
    Kyoto Encyclopedia of Genes and Genome Minoru Kanehisa Institute for Chemical Research, Kyoto University HFSPO Workshop, Strasbourg, November 18, 2016 The KEGG Databases Category Database Content PATHWAY KEGG pathway maps Systems information BRITE BRITE functional hierarchies MODULE KEGG modules KO (KEGG ORTHOLOGY) KO groups for functional orthologs Genomic information GENOME KEGG organisms, viruses and addendum GENES / SSDB Genes and proteins / sequence similarity COMPOUND Chemical compounds GLYCAN Glycans Chemical information REACTION / RCLASS Reactions / reaction classes ENZYME Enzyme nomenclature DISEASE Human diseases DRUG / DGROUP Drugs / drug groups Health information ENVIRON Health-related substances (KEGG MEDICUS) JAPIC Japanese drug labels DailyMed FDA drug labels 12 manually curated original DBs 3 DBs taken from outside sources and given original annotations (GENOME, GENES, ENZYME) 1 computationally generated DB (SSDB) 2 outside DBs (JAPIC, DailyMed) KEGG is widely used for functional interpretation and practical application of genome sequences and other high-throughput data KO PATHWAY GENOME BRITE DISEASE GENES MODULE DRUG Genome Molecular High-level Practical Metagenome functions functions applications Transcriptome etc. Metabolome Glycome etc. COMPOUND GLYCAN REACTION Funding Annual budget Period Funding source (USD) 1995-2010 Supported by 10+ grants from Ministry of Education, >2 M Japan Society for Promotion of Science (JSPS) and Japan Science and Technology Agency (JST) 2011-2013 Supported by National Bioscience Database Center 0.8 M (NBDC) of JST 2014-2016 Supported by NBDC 0.5 M 2017- ? 1995 KEGG website made freely available 1997 KEGG FTP site made freely available 2011 Plea to support KEGG KEGG FTP academic subscription introduced 1998 First commercial licensing Contingency Plan 1999 Pathway Solutions Inc.
    [Show full text]
  • 3 13437143.Pdf
    Title Integrative Annotation of 21,037 Human Genes Validated by Full-Length cDNA Clones Imanishi, Tadashi; Itoh, Takeshi; Suzuki, Yutaka; O'Donovan, Claire; Fukuchi, Satoshi; Koyanagi, Kanako O.; Barrero, Roberto A.; Tamura, Takuro; Yamaguchi-Kabata, Yumi; Tanino, Motohiko; Yura, Kei; Miyazaki, Satoru; Ikeo, Kazuho; Homma, Keiichi; Kasprzyk, Arek; Nishikawa, Tetsuo; Hirakawa, Mika; Thierry-Mieg, Jean; Thierry-Mieg, Danielle; Ashurst, Jennifer; Jia, Libin; Nakao, Mitsuteru; Thomas, Michael A.; Mulder, Nicola; Karavidopoulou, Youla; Jin, Lihua; Kim, Sangsoo; Yasuda, Tomohiro; Lenhard, Boris; Eveno, Eric; Suzuki, Yoshiyuki; Yamasaki, Chisato; Takeda, Jun-ichi; Gough, Craig; Hilton, Phillip; Fujii, Yasuyuki; Sakai, Hiroaki; Tanaka, Susumu; Amid, Clara; Bellgard, Matthew; Bonaldo, Maria de Fatima; Bono, Hidemasa; Bromberg, Susan K.; Brookes, Anthony J.; Bruford, Elspeth; Carninci, Piero; Chelala, Claude; Couillault, Christine; Souza, Sandro J. de; Debily, Marie-Anne; Devignes, Marie-Dominique; Dubchak, Inna; Endo, Toshinori; Estreicher, Anne; Eyras, Eduardo; Fukami-Kobayashi, Kaoru; R. Gopinath, Gopal; Graudens, Esther; Hahn, Yoonsoo; Han, Michael; Han, Ze-Guang; Hanada, Kousuke; Hanaoka, Hideki; Harada, Erimi; Hashimoto, Katsuyuki; Hinz, Ursula; Hirai, Momoki; Hishiki, Teruyoshi; Hopkinson, Ian; Imbeaud, Sandrine; Inoko, Hidetoshi; Kanapin, Alexander; Kaneko, Yayoi; Kasukawa, Takeya; Kelso, Janet; Kersey, Author(s) Paul; Kikuno, Reiko; Kimura, Kouichi; Korn, Bernhard; Kuryshev, Vladimir; Makalowska, Izabela; Makino, Takashi; Mano, Shuhei;
    [Show full text]
  • Open Data, Open Source, and Open Standards in Chemistry: the Blue Obelisk Five Years On" Journal of Cheminformatics Vol
    Oral Roberts University Digital Showcase College of Science and Engineering Faculty College of Science and Engineering Research and Scholarship 10-14-2011 Open Data, Open Source, and Open Standards in Chemistry: The lueB Obelisk five years on Andrew Lang Noel M. O'Boyle Rajarshi Guha National Institutes of Health Egon Willighagen Maastricht University Samuel Adams See next page for additional authors Follow this and additional works at: http://digitalshowcase.oru.edu/cose_pub Part of the Chemistry Commons Recommended Citation Andrew Lang, Noel M O'Boyle, Rajarshi Guha, Egon Willighagen, et al.. "Open Data, Open Source, and Open Standards in Chemistry: The Blue Obelisk five years on" Journal of Cheminformatics Vol. 3 Iss. 37 (2011) Available at: http://works.bepress.com/andrew-sid-lang/ 19/ This Article is brought to you for free and open access by the College of Science and Engineering at Digital Showcase. It has been accepted for inclusion in College of Science and Engineering Faculty Research and Scholarship by an authorized administrator of Digital Showcase. For more information, please contact [email protected]. Authors Andrew Lang, Noel M. O'Boyle, Rajarshi Guha, Egon Willighagen, Samuel Adams, Jonathan Alvarsson, Jean- Claude Bradley, Igor Filippov, Robert M. Hanson, Marcus D. Hanwell, Geoffrey R. Hutchison, Craig A. James, Nina Jeliazkova, Karol M. Langner, David C. Lonie, Daniel M. Lowe, Jerome Pansanel, Dmitry Pavlov, Ola Spjuth, Christoph Steinbeck, Adam L. Tenderholt, Kevin J. Theisen, and Peter Murray-Rust This article is available at Digital Showcase: http://digitalshowcase.oru.edu/cose_pub/34 Oral Roberts University From the SelectedWorks of Andrew Lang October 14, 2011 Open Data, Open Source, and Open Standards in Chemistry: The Blue Obelisk five years on Andrew Lang Noel M O'Boyle Rajarshi Guha, National Institutes of Health Egon Willighagen, Maastricht University Samuel Adams, et al.
    [Show full text]
  • Article Reference
    Article Incidences of problematic cell lines are lower in papers that use RRIDs to identify cell lines BABIC, Zeljana, et al. Abstract The use of misidentified and contaminated cell lines continues to be a problem in biomedical research. Research Resource Identifiers (RRIDs) should reduce the prevalence of misidentified and contaminated cell lines in the literature by alerting researchers to cell lines that are on the list of problematic cell lines, which is maintained by the International Cell Line Authentication Committee (ICLAC) and the Cellosaurus database. To test this assertion, we text-mined the methods sections of about two million papers in PubMed Central, identifying 305,161 unique cell-line names in 150,459 articles. We estimate that 8.6% of these cell lines were on the list of problematic cell lines, whereas only 3.3% of the cell lines in the 634 papers that included RRIDs were on the problematic list. This suggests that the use of RRIDs is associated with a lower reported use of problematic cell lines. Reference BABIC, Zeljana, et al. Incidences of problematic cell lines are lower in papers that use RRIDs to identify cell lines. eLife, 2019, vol. 8, p. e41676 DOI : 10.7554/eLife.41676 PMID : 30693867 Available at: http://archive-ouverte.unige.ch/unige:119832 Disclaimer: layout of this document may differ from the published version. 1 / 1 FEATURE ARTICLE META-RESEARCH Incidences of problematic cell lines are lower in papers that use RRIDs to identify cell lines Abstract The use of misidentified and contaminated cell lines continues to be a problem in biomedical research.
    [Show full text]
  • Open Data, Open Source and Open Standards in Chemistry: the Blue Obelisk five Years On
    Open Data, Open Source and Open Standards in chemistry: The Blue Obelisk ¯ve years on Noel M O'Boyle¤1 , Rajarshi Guha2 , Egon L Willighagen3 , Samuel E Adams4 , Jonathan Alvarsson5 , Richard L Apodaca6 , Jean-Claude Bradley7 , Igor V Filippov8 , Robert M Hanson9 , Marcus D Hanwell10 , Geo®rey R Hutchison11 , Craig A James12 , Nina Jeliazkova13 , Andrew SID Lang14 , Karol M Langner15 , David C Lonie16 , Daniel M Lowe4 , J¶er^omePansanel17 , Dmitry Pavlov18 , Ola Spjuth5 , Christoph Steinbeck19 , Adam L Tenderholt20 , Kevin J Theisen21 , Peter Murray-Rust4 1Analytical and Biological Chemistry Research Facility, Cavanagh Pharmacy Building, University College Cork, College Road, Cork, Co. Cork, Ireland 2NIH Center for Translational Therapeutics, 9800 Medical Center Drive, Rockville, MD 20878, USA 3Division of Molecular Toxicology, Institute of Environmental Medicine, Nobels vaeg 13, Karolinska Institutet, 171 77 Stockholm, Sweden 4Unilever Centre for Molecular Sciences Informatics, Department of Chemistry, University of Cambridge, Lens¯eld Road, CB2 1EW, UK 5Department of Pharmaceutical Biosciences, Uppsala University, Box 591, 751 24 Uppsala, Sweden 6Metamolecular, LLC, 8070 La Jolla Shores Drive #464, La Jolla, CA 92037, USA 7Department of Chemistry, Drexel University, 32nd and Chestnut streets, Philadelphia, PA 19104, USA 8Chemical Biology Laboratory, Basic Research Program, SAIC-Frederick, Inc., NCI-Frederick, Frederick, MD 21702, USA 9St. Olaf College, 1520 St. Olaf Ave., North¯eld, MN 55057, USA 10Kitware, Inc., 28 Corporate Drive, Clifton Park, NY 12065, USA 11Department of Chemistry, University of Pittsburgh, 219 Parkman Avenue, Pittsburgh, PA 15260, USA 12eMolecules Inc., 380 Stevens Ave., Solana Beach, California 92075, USA 13Ideaconsult Ltd., 4.A.Kanchev str., So¯a 1000, Bulgaria 14Department of Engineering, Computer Science, Physics, and Mathematics, Oral Roberts University, 7777 S.
    [Show full text]
  • Genbank by Walter B
    GenBank by Walter B. Goad o understand the significance of the information stored in memory—DNA—to the stuff of activity—proteins. The idea that GenBank, you need to know a little about molecular somehow the bases in DNA determine the amino acids in proteins genetics. What that field deals with is self-replication—the had been around for some time. In fact, George Gamow suggested in T process unique to life—and mutation and 1954, after learning about the structure proposed for DNA, that a recombination—the processes responsible for evolution-at the triplet of bases corresponded to an amino acid. That suggestion was fundamental level of the genes in DNA. This approach of working shown to be true, and by 1965 most of the genetic code had been from the blueprint, so to speak, of a living system is very powerful, deciphered. Also worked out in the ’60s were many details of what and studies of many other aspects of life—the process of learning, for Crick called the central dogma of molecular genetics—the now example—are now utilizing molecular genetics. firmly established fact that DNA is not translated directly to proteins Molecular genetics began in the early ’40s and was at first but is first transcribed to messenger RNA. This molecule, a nucleic controversial because many of the people involved had been trained acid like DNA, then serves as the template for protein synthesis. in the physical sciences rather than the biological sciences, and yet These great advances prompted a very distinguished molecular they were answering questions that biologists had been asking for geneticist to predict, in 1969, that biology was just about to end since years.
    [Show full text]
  • Integrative Annotation of 21,037 Human Genes Validated by Full-Length Cdna Clones
    Title Integrative Annotation of 21,037 Human Genes Validated by Full-Length cDNA Clones Imanishi, Tadashi; Itoh, Takeshi; Suzuki, Yutaka; O'Donovan, Claire; Fukuchi, Satoshi; Koyanagi, Kanako O.; Barrero, Roberto A.; Tamura, Takuro; Yamaguchi-Kabata, Yumi; Tanino, Motohiko; Yura, Kei; Miyazaki, Satoru; Ikeo, Kazuho; Homma, Keiichi; Kasprzyk, Arek; Nishikawa, Tetsuo; Hirakawa, Mika; Thierry-Mieg, Jean; Thierry-Mieg, Danielle; Ashurst, Jennifer; Jia, Libin; Nakao, Mitsuteru; Thomas, Michael A.; Mulder, Nicola; Karavidopoulou, Youla; Jin, Lihua; Kim, Sangsoo; Yasuda, Tomohiro; Lenhard, Boris; Eveno, Eric; Suzuki, Yoshiyuki; Yamasaki, Chisato; Takeda, Jun-ichi; Gough, Craig; Hilton, Phillip; Fujii, Yasuyuki; Sakai, Hiroaki; Tanaka, Susumu; Amid, Clara; Bellgard, Matthew; Bonaldo, Maria de Fatima; Bono, Hidemasa; Bromberg, Susan K.; Brookes, Anthony J.; Bruford, Elspeth; Carninci, Piero; Chelala, Claude; Couillault, Christine; Souza, Sandro J. de; Debily, Marie-Anne; Devignes, Marie-Dominique; Dubchak, Inna; Endo, Toshinori; Estreicher, Anne; Eyras, Eduardo; Fukami-Kobayashi, Kaoru; R. Gopinath, Gopal; Graudens, Esther; Hahn, Yoonsoo; Han, Michael; Han, Ze-Guang; Hanada, Kousuke; Hanaoka, Hideki; Harada, Erimi; Hashimoto, Katsuyuki; Hinz, Ursula; Hirai, Momoki; Hishiki, Teruyoshi; Hopkinson, Ian; Imbeaud, Sandrine; Inoko, Hidetoshi; Kanapin, Alexander; Kaneko, Yayoi; Kasukawa, Takeya; Kelso, Janet; Kersey, Author(s) Paul; Kikuno, Reiko; Kimura, Kouichi; Korn, Bernhard; Kuryshev, Vladimir; Makalowska, Izabela; Makino, Takashi; Mano, Shuhei;
    [Show full text]
  • General Assembly and Consortium Meeting 2020
    EJP RD GA and Consortium meeting Final program EJP RD General Assembly and Consortium meeting 2020 Online September 14th – 18th Page 1 of 46 EJP RD GA and Consortium meeting Final program Program at a glance: Page 2 of 46 EJP RD GA and Consortium meeting Final program Table of contents Program per day ....................................................................................................................... 4 Plenary GA Session .................................................................................................................. 8 Parallel Session Pillar 1 .......................................................................................................... 10 Parallel Session Pillar 2 .......................................................................................................... 11 Parallel Session Pillar 3 .......................................................................................................... 14 Parallel Session Pillar 4 .......................................................................................................... 16 Experimental data resources available in the EJP RD ............................................................. 18 Introduction to translational research for clinical or pre-clinical researchers .......................... 21 Funding opportunities in EJP RD & How to build a successful proposal ............................... 22 RD-Connect Genome-Phenome Analysis Platform ................................................................. 26 Improvement
    [Show full text]
  • 2003 Mulder Nucl Acids Res {22
    The InterPro Database, 2003 brings increased coverage and new features Nicola J Mulder, Rolf Apweiler, Teresa K Attwood, Amos Bairoch, Daniel Barrell, Alex Bateman, David Binns, Margaret Biswas, Paul Bradley, Peer Bork, et al. To cite this version: Nicola J Mulder, Rolf Apweiler, Teresa K Attwood, Amos Bairoch, Daniel Barrell, et al.. The InterPro Database, 2003 brings increased coverage and new features. Nucleic Acids Research, Oxford University Press, 2003, 31 (1), pp.315-318. 10.1093/nar/gkg046. hal-01214149 HAL Id: hal-01214149 https://hal.archives-ouvertes.fr/hal-01214149 Submitted on 9 Oct 2015 HAL is a multi-disciplinary open access L’archive ouverte pluridisciplinaire HAL, est archive for the deposit and dissemination of sci- destinée au dépôt et à la diffusion de documents entific research documents, whether they are pub- scientifiques de niveau recherche, publiés ou non, lished or not. The documents may come from émanant des établissements d’enseignement et de teaching and research institutions in France or recherche français ou étrangers, des laboratoires abroad, or from public or private research centers. publics ou privés. # 2003 Oxford University Press Nucleic Acids Research, 2003, Vol. 31, No. 1 315–318 DOI: 10.1093/nar/gkg046 The InterPro Database, 2003 brings increased coverage and new features Nicola J. Mulder1,*, Rolf Apweiler1, Teresa K. Attwood3, Amos Bairoch4, Daniel Barrell1, Alex Bateman2, David Binns1, Margaret Biswas5, Paul Bradley1,3, Peer Bork6, Phillip Bucher7, Richard R. Copley8, Emmanuel Courcelle9, Ujjwal Das1, Richard Durbin2, Laurent Falquet7, Wolfgang Fleischmann1, Sam Griffiths-Jones2, Downloaded from Daniel Haft10, Nicola Harte1, Nicolas Hulo4, Daniel Kahn9, Alexander Kanapin1, Maria Krestyaninova1, Rodrigo Lopez1, Ivica Letunic6, David Lonsdale1, Ville Silventoinen1, Sandra E.
    [Show full text]
  • ACCEPTED PAPERS – ABSTRACTS (As of May 5, 2006) Comparative
    ACCEPTED PAPERS – ABSTRACTS (as of May 5, 2006) Comparative Genomics Hairpins in a haystack: recognizing microrna precursors in comparative genomics data Author(s): Jana Hertel, Peter F. Stadler Recently, genome wide surveys for non-coding RNAs have provided evidence for tens of thousands of previously undescribed evolutionary conserved RNAs with distinctive secondary structures. The annotation of these putative ncRNAs, however, remains a difficult problem. Here we describe an SVM-based approach that, in conjunction with a non-stringent filter for consensus secondary structures, is capable of efficiently recognizing microRNA precursors in multiple sequence alignments. The software was applied to recent genome-wide RNAz surveys of mammals, urochordates, and nematodes. Keywords: miRNA, support vector machine, non-coding RNA Comparative Genomics Comparative genomics reveals unusually long motifs in mammalian genomes Author(s): Neil Jones, Pavel Pevzner Motivation: The recent discovery of the first small modulatory RNA (smRNA) presents the challenge of finding other molecules of similar length and conservation level. Unlike short interfering RNA (siRNA) and micro-RNA (miRNA), effective computational and experimental screening methods are not currently known for this species of RNA molecule, and the discovery of the one known example was partly fortuitous because it happened to be complementary to a well-studied DNA binding motif (the Neuron Restrictive Silencer Element). Results: The existing comparative genomics approaches (e.g., phylogenetic footprinting) rely on alignments of orthologous regions across multiple genomes. This approach, while extremely valuable, is not suitable for finding motifs with highly diverged ``non-alignable'' flanking regions. Here we show that several unusually long and well conserved motifs can be discovered de novo through a comparative genomics approach that does not require an alignment of orthologous upstream regions.
    [Show full text]
  • MPGM: Scalable and Accurate Multiple Network Alignment
    MPGM: Scalable and Accurate Multiple Network Alignment Ehsan Kazemi1 and Matthias Grossglauser2 1Yale Institute for Network Science, Yale University 2School of Computer and Communication Sciences, EPFL Abstract Protein-protein interaction (PPI) network alignment is a canonical operation to transfer biological knowledge among species. The alignment of PPI-networks has many applica- tions, such as the prediction of protein function, detection of conserved network motifs, and the reconstruction of species’ phylogenetic relationships. A good multiple-network align- ment (MNA), by considering the data related to several species, provides a deep understand- ing of biological networks and system-level cellular processes. With the massive amounts of available PPI data and the increasing number of known PPI networks, the problem of MNA is gaining more attention in the systems-biology studies. In this paper, we introduce a new scalable and accurate algorithm, called MPGM, for aligning multiple networks. The MPGM algorithm has two main steps: (i) SEEDGENERA- TION and (ii) MULTIPLEPERCOLATION. In the first step, to generate an initial set of seed tuples, the SEEDGENERATION algorithm uses only protein sequence similarities. In the second step, to align remaining unmatched nodes, the MULTIPLEPERCOLATION algorithm uses network structures and the seed tuples generated from the first step. We show that, with respect to different evaluation criteria, MPGM outperforms the other state-of-the-art algo- rithms. In addition, we guarantee the performance of MPGM under certain classes of net- work models. We introduce a sampling-based stochastic model for generating k correlated networks. We prove that for this model if a sufficient number of seed tuples are available, the MULTIPLEPERCOLATION algorithm correctly aligns almost all the nodes.
    [Show full text]