Ensembl Genomes 2020––Enabling Non-Vertebrate Genomic Research Kevin L

Total Page:16

File Type:pdf, Size:1020Kb

Load more

Published online 10 October 2019 Nucleic Acids Research, 2020, Vol. 48, Database issue D689–D695 doi: 10.1093/nar/gkz890 Ensembl Genomes 2020––enabling non-vertebrate genomic research Kevin L. Howe 1, Bruno Contreras-Moreira1, Nishadi De Silva1, Gareth Maslen1, Wasiu Akanni1, James Allen1, Jorge Alvarez-Jarreta1, Matthieu Barba1, Dan M. Bolser 1, Lahcen Cambell1, Manuel Carbajo1, Marc Chakiachvili1, Mikkel Christensen1, Carla Cummins1, Alayne Cuzick2, Paul Davis1, Silvie Fexova1, Astrid Gall1, Nancy George1, Laurent Gil1, Parul Gupta3, Kim E. Hammond-Kosack 2, Erin Haskell1, Sarah E. Hunt 1, Pankaj Jaiswal 3, Sophie H. Janacek1, Paul J. Kersey1, Nick Langridge1, Uma Maheswari1, Thomas Maurel1, Mark D. McDowall1, Ben Moore1, Matthieu Muffato1, Guy Naamati1, Sushma Naithani 3, Andrew Olson4, Irene Papatheodorou1, Mateus Patricio1, Michael Paulini1, Helder Pedro 1,EmilyPerry 1, Justin Preece3, Marc Rosello1, Matthew Russell 1, Vasily Sitnik1, Daniel M. Staines1, Joshua Stein4, Marcela K. Tello-Ruiz4, Stephen J. Trevanion1, Martin Urban2, Sharon Wei4, Doreen Ware4,5, Gary Williams1, Andrew D. Yates 1 and Paul Flicek 1,* 1European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, UK, 2Department of Biointeractions and Crop Protection, Rothamsted Research, Harpenden, Hertfordshire AL5 2JQ, UK, 3Department of Botany and Plant Pathology, Oregon State University, Corvallis, OR 97331, USA, 4Cold Spring Harbor Laboratory, 1 Bungtown Rd, Cold Spring Harbor, NY 11724, USA and 5USDA ARS NAA Robert W. Holley Center for Agriculture and Health, Agricultural Research Service, Ithaca, NY 14853, USA Received September 14, 2019; Revised September 29, 2019; Editorial Decision September 30, 2019; Accepted October 02, 2019 ABSTRACT Ensembl project, which forms a key part of our fu- ture strategy for dealing with the increasing quantity Ensembl Genomes (http://www.ensemblgenomes. of available genome-scale data across the tree of life. org) is an integrating resource for genome-scale data from non-vertebrate species, complementing the re- sources for vertebrate genomics developed in the INTRODUCTION context of the Ensembl project (http://www.ensembl. The goal of Ensembl Genomes is to support and facili- org). Together, the two resources provide a consis- tate non-vertebrate genome research by integrating high- tent set of interfaces to genomic data across the quality reference genomes and annotation for every species tree of life, including reference genome sequence, for which this is available, providing tools and displays gene models, transcriptional data, genetic variation that allow users to interrogate and compare these genomes, and integrating with and connecting to other resources and comparative analysis. Data may be accessed providing non-vertebrate reference genomic data. To this via our website, online tools platform and program- end, we work closely with the Ensembl project (1)topro- matic interfaces, with updates made four times per duce five additional sites, each focused on a particular non- year (in synchrony with Ensembl). Here, we provide vertebrate clade of life: bacteria (http://bacteria.ensembl. an overview of Ensembl Genomes, with a focus on org), protists (http://protists.ensembl.org), fungi (http:// recent developments. These include the continued fungi.ensembl.org), plants (http://plants.ensembl.org)and growth, more robust and reproducible sets of ortho- invertebrate metazoa (http://metazoa.ensembl.org). These logues and paralogues, and enriched views of gene sites are updated four times a year, using the same platform expression and gene function in plants. Finally, we as, and in synchrony with, Ensembl. report on our continued deeper integration with the Core data available for all species include genome se- quence and annotations of protein-coding and non-coding *To whom correspondence should be addressed. Tel: +44 1223 494468; Email: [email protected] C The Author(s) 2019. Published by Oxford University Press on behalf of Nucleic Acids Research. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited. D690 Nucleic Acids Research, 2020, Vol. 48, Database issue genes. Transcriptional data, genetic variation and com- tinued access to the website beyond when they have been parative analysis data are additionally available for many superseded by new releases. Archive releases to date are species. For most species, data are imported directly from 37 (e.g. http://eg37-plants.ensembl.org) and 40 (e.g. http: the archives of the International Nucleotide Sequence //eg40-fungi.ensembl.org). Database Collaboration (INSDC) (2), the European Vari- Finally, for programmatic access, we provide a compre- ation Archive (http://www.ebi.ac.uk/eva) and other open- hensive language-agnostic RESTful Application Program- access sources. For a few species of particular research or ming Interface (API). A significant update this year has socioeconomic importance, additional data sets are identi- been the consolidation of the APIs for Ensembl and En- fied and integrated. However, according to our agreement sembl Genomes into a single service (https://rest.ensembl. with NCBI and UCSC, we only present data on genome as- org), for the first time providing access to both vertebrate semblies that have been accessioned by the INSDC. and non-vertebrate data through a single programmatic in- We collaborate closely with a number of other inter- terface. A training manual for the consolidated API can national resources in various domains of life, including be found at https://www.ebi.ac.uk/training/online/course/ Gramene (http://www.gramene.org)(3) for plants, Vector- ensembl-rest-api. Base (http://www.vectorbase.org)(4) for invertebrate vec- tors of human pathogens and WormBase (http://www. NEW AND IMPROVED GENOMES wormbase.org)(5) for nematodes and flatworms. Together, we develop common data sets and integration/analysis The past 2 years have seen continued growth in our genome workflows, which are made available through both Ensembl collections (see Table 1). Contents of Ensembl Bacteria have Genomes and project-specific websites and services. been frozen while our focus has been on other divisions in this period. We are working with partners to develop a co- herent strategy for prokaryotes, especially in the context of ACCESS the growing availability of metagenomic data, and antici- All data in the resource are open access and can be ex- pate significant progress on this in the next year. plored in multiple ways. Interactive and visual access is pro- Of the 250 genomes added in Ensembl Fungi and En- vided through our websites. A key component of these sites sembl Protists, significant additions include Puccinia coro- is our embedded genome browser, enabling users to scroll nata, a new and fast-spreading pathogen of barley and oats; through a graphical representation of a DNA molecule at Phytophthora megakarya, an oomycete pathogen causing different levels of resolution and visualize various types of black pod disease in cocoa trees; and the ice-loving diatom annotations as tracks on the browser, including gene mod- Fragilariopsis cylindrus, found in Arctic and Antarctic sea- els, genomic variants, repetitive elements and homology se- water. quences aligned to the genome. We also display functional In metazoa, notable additions have been Folsomia can- information on genes and gene products, imported from dida and Orchesella cincta as examples of organisms impor- the UniProt Knowledgebase (6) and analysis of peptide se- tant for the assessment of soil quality; Bombus terrestris (the quence (using the classification tool InterProScan (7)). Var- buff-tailed bumblebee), a declining insect species; Daph- ious tools for text and sequence search, data upload and nia magna, which is used in studying toxicity (such as data analysis are provided, allowing users to examine their polystyrene nanoplastics, lead and pesticides) in aquatic en- own data in the context of the reference sequence and an- vironments; Teleopsis dalmanni (the stalk eyed fly), a model notation. organism for studying sexual selection mechanisms; and All data are made available as database dumps, and com- the disease vectors Leptotrombidium deliense (scrub typhus) mon data sets (e.g. DNA, RNA and protein sequence sets, and Culicoides sonorensis (bluetongue arbovirus in live- and sequence alignments) can be directly downloaded in stock). bulk via FTP (ftp://ftp.ensemblgenomes.org)inavarietyof In plants, we have significantly increased coverage of formats. For certain data types, the web application utilizes asterids, including Actinidia chinensis (kiwifruit), Coffea data files stored in archival resources (such as the European canephora (robusta coffee), Cynara cardunculus (globe arti- Nucleotide Archive), avoiding the need for database builds choke), Daucus carota (carrot), Capsicum annuum (hot pep- and improving the speed of response. These files are also per) and Helianthus annuus (sunflower). Elsewhere in the made available to users. taxonomy, we have included Arabidopsis halleri, a brassica Data are also made available through a series of data model for ecological genomics due to its heavy metal hyper- warehouses, optimized around common (gene- and variant- accumulation ability; and Marchantia polymorpha,aliver-
Recommended publications
  • Ensembl Genomes: Extending Ensembl Across the Taxonomic Space P

    Ensembl Genomes: Extending Ensembl Across the Taxonomic Space P

    Published online 1 November 2009 Nucleic Acids Research, 2010, Vol. 38, Database issue D563–D569 doi:10.1093/nar/gkp871 Ensembl Genomes: Extending Ensembl across the taxonomic space P. J. Kersey*, D. Lawson, E. Birney, P. S. Derwent, M. Haimel, J. Herrero, S. Keenan, A. Kerhornou, G. Koscielny, A. Ka¨ ha¨ ri, R. J. Kinsella, E. Kulesha, U. Maheswari, K. Megy, M. Nuhn, G. Proctor, D. Staines, F. Valentin, A. J. Vilella and A. Yates EMBL-European Bioinformatics Institute, Wellcome Trust Genome Campus, Cambridge CB10 1SD, UK Received August 14, 2009; Revised September 28, 2009; Accepted September 29, 2009 ABSTRACT nucleotide archives; numerous other genomes exist in states of partial assembly and annotation; thousands of Ensembl Genomes (http://www.ensemblgenomes viral genomes sequences have also been generated. .org) is a new portal offering integrated access to Moreover, the increasing use of high-throughput genome-scale data from non-vertebrate species sequencing technologies is rapidly reducing the cost of of scientific interest, developed using the Ensembl genome sequencing, leading to an accelerating rate of genome annotation and visualisation platform. data production. This not only makes it likely that in Ensembl Genomes consists of five sub-portals (for the near future, the genomes of all species of scientific bacteria, protists, fungi, plants and invertebrate interest will be sequenced; but also the genomes of many metazoa) designed to complement the availability individuals, with the possibility of providing accurate and of vertebrate genomes in Ensembl. Many of the sophisticated annotation through the similarly low-cost databases supporting the portal have been built in application of functional assays.
  • Abstracts Genome 10K & Genome Science 29 Aug - 1 Sept 2017 Norwich Research Park, Norwich, Uk

    Abstracts Genome 10K & Genome Science 29 Aug - 1 Sept 2017 Norwich Research Park, Norwich, Uk

    Genome 10K c ABSTRACTS GENOME 10K & GENOME SCIENCE 29 AUG - 1 SEPT 2017 NORWICH RESEARCH PARK, NORWICH, UK Genome 10K c 48 KEYNOTE SPEAKERS ............................................................................................................................... 1 Dr Adam Phillippy: Towards the gapless assembly of complete vertebrate genomes .................... 1 Prof Kathy Belov: Saving the Tasmanian devil from extinction ......................................................... 1 Prof Peter Holland: Homeobox genes and animal evolution: from duplication to divergence ........ 2 Dr Hilary Burton: Genomics in healthcare: the challenges of complexity .......................................... 2 INVITED SPEAKERS ................................................................................................................................. 3 Vertebrate Genomics ........................................................................................................................... 3 Alex Cagan: Comparative genomics of animal domestication .......................................................... 3 Plant Genomics .................................................................................................................................... 4 Ksenia Krasileva: Evolution of plant Immune receptors ..................................................................... 4 Andrea Harper: Using Associative Transcriptomics to predict tolerance to ash dieback disease in European ash trees ............................................................................................................
  • The ELIXIR Core Data Resources: ​Fundamental Infrastructure for The

    The ELIXIR Core Data Resources: ​Fundamental Infrastructure for The

    Supplementary Data: The ELIXIR Core Data Resources: fundamental infrastructure ​ for the life sciences The “Supporting Material” referred to within this Supplementary Data can be found in the Supporting.Material.CDR.infrastructure file, DOI: 10.5281/zenodo.2625247 (https://zenodo.org/record/2625247). ​ ​ Figure 1. Scale of the Core Data Resources Table S1. Data from which Figure 1 is derived: Year 2013 2014 2015 2016 2017 Data entries 765881651 997794559 1726529931 1853429002 2715599247 Monthly user/IP addresses 1700660 2109586 2413724 2502617 2867265 FTEs 270 292.65 295.65 289.7 311.2 Figure 1 includes data from the following Core Data Resources: ArrayExpress, BRENDA, CATH, ChEBI, ChEMBL, EGA, ENA, Ensembl, Ensembl Genomes, EuropePMC, HPA, IntAct /MINT , InterPro, PDBe, PRIDE, SILVA, STRING, UniProt ● Note that Ensembl’s compute infrastructure physically relocated in 2016, so “Users/IP address” data are not available for that year. In this case, the 2015 numbers were rolled forward to 2016. ● Note that STRING makes only minor releases in 2014 and 2016, in that the interactions are re-computed, but the number of “Data entries” remains unchanged. The major releases that change the number of “Data entries” happened in 2013 and 2015. So, for “Data entries” , the number for 2013 was rolled forward to 2014, and the number for 2015 was rolled forward to 2016. The ELIXIR Core Data Resources: fundamental infrastructure for the life sciences ​ 1 Figure 2: Usage of Core Data Resources in research The following steps were taken: 1. API calls were run on open access full text articles in Europe PMC to identify articles that ​ ​ mention Core Data Resource by name or include specific data record accession numbers.
  • Annual Scientific Report 2013 on the Cover Structure 3Fof in the Protein Data Bank, Determined by Laponogov, I

    Annual Scientific Report 2013 on the Cover Structure 3Fof in the Protein Data Bank, Determined by Laponogov, I

    EMBL-European Bioinformatics Institute Annual Scientific Report 2013 On the cover Structure 3fof in the Protein Data Bank, determined by Laponogov, I. et al. (2009) Structural insight into the quinolone-DNA cleavage complex of type IIA topoisomerases. Nature Structural & Molecular Biology 16, 667-669. © 2014 European Molecular Biology Laboratory This publication was produced by the External Relations team at the European Bioinformatics Institute (EMBL-EBI) A digital version of the brochure can be found at www.ebi.ac.uk/about/brochures For more information about EMBL-EBI please contact: [email protected] Contents Introduction & overview 3 Services 8 Genes, genomes and variation 8 Molecular atlas 12 Proteins and protein families 14 Molecular and cellular structures 18 Chemical biology 20 Molecular systems 22 Cross-domain tools and resources 24 Research 26 Support 32 ELIXIR 36 Facts and figures 38 Funding & resource allocation 38 Growth of core resources 40 Collaborations 42 Our staff in 2013 44 Scientific advisory committees 46 Major database collaborations 50 Publications 52 Organisation of EMBL-EBI leadership 61 2013 EMBL-EBI Annual Scientific Report 1 Foreword Welcome to EMBL-EBI’s 2013 Annual Scientific Report. Here we look back on our major achievements during the year, reflecting on the delivery of our world-class services, research, training, industry collaboration and European coordination of life-science data. The past year has been one full of exciting changes, both scientifically and organisationally. We unveiled a new website that helps users explore our resources more seamlessly, saw the publication of ground-breaking work in data storage and synthetic biology, joined the global alliance for global health, built important new relationships with our partners in industry and celebrated the launch of ELIXIR.
  • Strategic Plan 2011-2016

    Strategic Plan 2011-2016

    Strategic Plan 2011-2016 Wellcome Trust Sanger Institute Strategic Plan 2011-2016 Mission The Wellcome Trust Sanger Institute uses genome sequences to advance understanding of the biology of humans and pathogens in order to improve human health. -i- Wellcome Trust Sanger Institute Strategic Plan 2011-2016 - ii - Wellcome Trust Sanger Institute Strategic Plan 2011-2016 CONTENTS Foreword ....................................................................................................................................1 Overview .....................................................................................................................................2 1. History and philosophy ............................................................................................................ 5 2. Organisation of the science ..................................................................................................... 5 3. Developments in the scientific portfolio ................................................................................... 7 4. Summary of the Scientific Programmes 2011 – 2016 .............................................................. 8 4.1 Cancer Genetics and Genomics ................................................................................ 8 4.2 Human Genetics ...................................................................................................... 10 4.3 Pathogen Variation .................................................................................................. 13 4.4 Malaria
  • Browsing Genomes with Ensembl Annotation

    Browsing Genomes with Ensembl Annotation

    Browsing genomes with EnsEMBL Annotation • During recent years release of large amounts of sequence data • Raw sequence data are not so useful on its own. They are most valuable when provided with comprehensive good quality annotation CCCAACAAGAATGTAAAATCTTTAAGTGCCTGTTTTCATACTTATTTGACCACCCTATCTCTAGAATCTTGCATGATG TCTAGCCCTAGTAGGATCAAAAAATACTTACAAAGCAACTGAATAGCTACATGAATAGATGGATGAATAAATGCATG GGTGGATGGATGGATTAATGAAATCATTTATATGACTTAAAGTTTGCAGAGGAGTATCATATTTGGAAGGCAGTAAG GAAGTCTGTGTAGTCGATGGTAAAGGCAATTGGGAAGTTTGTTAGGCACAATAGGTCAAAATTTGTTTTTGAAGTCC TGTTACTTCACGTTTCTTTGTTTCACTTTCTTAAAACAGGAAACTCTTTTCTATGATCATTCTTCCAGGGCCTGGCTCT TCATCTGCAACCCAGTAATATCCCTAATGTCAAAAAGCTACTGGTTTAATTCGTGCCATTTTCAAAGAGGACTACTGA ATTCTGATGTGGCTTCAAACATTTAGGTTAGGCATATCTAATGGAGAACTTGCAGCCACACTGACTTGTAGTGAAAT ATCTATTTTGAGCCTGCCCAGTGTTGCTTAAATTGTAGTTTTCCTTGCCAGCTATTCATACAAGAGATGTGAGAAGCA CCATAAAAGGCGTTGTGAGGAGTTGTGGGGGAGTGAGGGAGAGAAGAGGTTGAAAAGCTTATTAGCTGCTGTACGG TAAAAGTGAGCTCTTACGGGAATGGGAATGTAGTTTTAGCCCTCCAGGGATTCTATTTAGCCCGCCAGGAATTAACC TTGACTATAAATAGGCCATCAATGACCTTTCCAGAGAATGTTCAGAGACCTCAACTTTGTTTAGAGATCTTGTGTGGG TGGAACTTCCTGTTTGCACACAGAGCAGCATAAAGCCCAGTTGCTTTGGGAAGTGTTTGGGACCAGATGGATTGTAG GGAGTAGGGTACAATACAGTCTGTTCTCCTCCAGCTCCTTCTTTCTGCAACATGGGGAAGAACAAACTCCTTCATCC AAGTCTGGTTCTTCTCCTCTTGGTCCTCCTGCCCACAGACGCCTCAGTCTCTGGAAAACCGTGAGTTCCACACAGAG AGCGTGAAGCATGAACCTAGAGTCCTTCATTTATTGCAGATTTTTCTTTATATCATTCCTTTTTCTTTCCTATGATACT GTCATCTTCTTATCTCTAAGATTCCTTCCAGATTTTACAAATCTAGTTTACTCATTACTTGCTTACTTTTAATCATTCT TCCCCAACTCTCTGAAGCTCTAATATGCAAAGCCTTCCTAAGGGGTGTCAGAAATTTTTAGCTTTTTAAAAGAATAAA
  • The Genomic Basis of Circadian and Circalunar Timing Adaptations in a Midge Tobias S

    The Genomic Basis of Circadian and Circalunar Timing Adaptations in a Midge Tobias S

    OPEN ARTICLE doi:10.1038/nature20151 The genomic basis of circadian and circalunar timing adaptations in a midge Tobias S. Kaiser1,2,3†, Birgit Poehn1,3, David Szkiba2, Marco Preussner4, Fritz J. Sedlazeck2†, Alexander Zrim2, Tobias Neumann1,2, Lam-Tung Nguyen2,5, Andrea J. Betancourt6, Thomas Hummel3,7, Heiko Vogel8, Silke Dorner1, Florian Heyd4, Arndt von Haeseler2,3,5 & Kristin Tessmar-Raible1,3 Organisms use endogenous clocks to anticipate regular environmental cycles, such as days and tides. Natural variants resulting in differently timed behaviour or physiology, known as chronotypes in humans, have not been well characterized at the molecular level. We sequenced the genome of Clunio marinus, a marine midge whose reproduction is timed by circadian and circalunar clocks. Midges from different locations show strain-specific genetic timing adaptations. We examined genetic variation in five C. marinus strains from different locations and mapped quantitative trait loci for circalunar and circadian chronotypes. The region most strongly associated with circadian chronotypes generates strain-specific differences in the abundance of calcium/calmodulin-dependent kinase II.1 (CaMKII.1) splice variants. As equivalent variants were shown to alter CaMKII activity in Drosophila melanogaster, and C. marinus (Cma)-CaMKII.1 increases the transcriptional activity of the dimer of the circadian proteins Cma-CLOCK and Cma-CYCLE, we suggest that modulation of alternative splicing is a mechanism for natural adaptation in circadian timing. Around the new or full moon, during a few specific hours surround- Our study aimed to identify the genetic basis of C. marinus adaptation ing low tide, millions of non-biting midges of the species C.
  • ALEXA: a Microarray Design Platform for Alternative Expression Analysis

    ALEXA: a Microarray Design Platform for Alternative Expression Analysis

    CORRESPONDENCE expression of alternative mRNA isoforms in 5-fluorouracil (5-FU)- ALEXA: a microarray design platform for sensitive and resistant colorectal cancer cell lines5 and compared alternative expression analysis the results to those from the Affymetrix ‘GeneChip Human Exon 1.0 ST’ array (see Supplementary Results, Supplementary Fig. 2 To the editor: Eukaryotic genomes are predicted to contain about and Supplementary Table 2 online). Genes and exons differentially 7,000–29,000 genes1. Each of these genes may be alternatively expressed between 5-FU–sensitive and resistant cells were identi- processed to produce multiple distinct mRNAs by alternative fied by both platforms (with significant overlap), but ALEXA arrays transcript initiation, splicing and polyadenylation (collectively provided additional information on the connectivity and boundar- referred to as alternative expression). Although analysis of avail- ies of exons (Table 1). Furthermore, alternative expression events able transcript resources indicates that up to ~75% of genes are identified by ALEXA were significantly enriched for known alterna- alternatively processed, most microarray expression platforms tive expression events represented in publicly available mRNA and cannot detect alternative transcripts2. expressed sequence tag (EST) databases (Supplementary Results methods Proof-of-principle experiments have described the use of oli- and Supplementary Data 1 online). Finally, we demonstrated the gonucleotide microarrays to profile transcript isoforms gener- advantage of the ALEXA approach by identifying several differen- ated by alternative expression, but resources to create such arrays tially expressed known and predicted isoforms with potential rele- are lacking3,4. To address this limitation we created a microarray vance to 5-FU resistance (Supplementary Fig. 3 and Supplementary .com/nature e design platform for alternative expression analysis (ALEXA), Tables 3 and 4 online).
  • Gene3d: Multi-Domain Annotations for Protein Sequence and Comparative Genome Analysis Jonathan G

    Gene3d: Multi-Domain Annotations for Protein Sequence and Comparative Genome Analysis Jonathan G

    D240–D245 Nucleic Acids Research, 2014, Vol. 42, Database issue Published online 21 November 2013 doi:10.1093/nar/gkt1205 Gene3D: Multi-domain annotations for protein sequence and comparative genome analysis Jonathan G. Lees1,*, David Lee1, Romain A. Studer1, Natalie L. Dawson1, Ian Sillitoe1, Sayoni Das1, Corin Yeats2, Benoit H. Dessailly1, Robert Rentzsch3 and Christine A. Orengo1 1Division of Biosciences, Institute of Structural and Molecular Biology, University College London, Gower Street, London WC1E 6BT, UK, 2Department of Infectious Disease Epidemiology, Imperial College London, St Mary’s Campus, Norfolk Place, London W2 1PG, UK and 3Robert Koch Institut, Research Group Bioinformatics Ng4, Nordufer 20, 13353 Berlin, Germany Received October 1, 2013; Revised November 2, 2013; Accepted November 4, 2013 Downloaded from ABSTRACT contain long disordered sections that are also of functional importance. The CATH database which focuses on the Gene3D (http://gene3d.biochem.ucl.ac.uk) is a data- folded domain, classifies structures in the PDB into their base of protein domain structure annotations for constituent domains, and each domain is subsequently http://nar.oxfordjournals.org/ protein sequences. Domains are predicted using a assigned to a single superfamily by homology (1,2). The library of profile HMMs from 2738 CATH super- classified domain structure sequences are used in a families. Gene3D assigns domain annotations to pipeline to build domain superfamily specific HMMs that Ensembl and UniProt sequence sets including are then used to identify domains in structurally >6000 cellular genomes and >20 million unique uncharacterized protein sequences. As some CATH protein sequences. This represents an increase of superfamilies can be large and functionally diverse, we 45% in the number of protein sequences since our recently introduced protocols for subdividing each super- last publication.
  • What's New in Rnacentral and Rfam

    What's New in Rnacentral and Rfam

    What’s new in RNAcentral and Rfam Anton Petrov [email protected] Benasque - July 20, 2018 What is https://www.vecteezy.com/vector-art/92726-question-mark-background-vector The non-coding RNA sequence database rnacentral.org • >10 million sequences • 27 databases • 800,000 species RNAcentral has lots of useful data • Sequence • Description • RNA type • Links to other databases • Genome locations • Publications • RNA modifications from MODOMICS and PDB http://rnacentral.org/rna/URS00005A4DCF/9606 Until recently just data aggregation, now additional analysis Two important new features 1. Quality control using Rfam 2. Comprehensive genome models mapping 1. Rfam models are used to annotate RNAcentral • ~90% of RNAcentral sequences match Rfam models • about 2% of RNAcentral sequences can be used to build new Rfam models Rfam annotations help detect: • truncated sequences • potential contamination • missing annotations 2. Comprehensive genome mapping for >250 species before after • genome mapping Ensembl genomes and blat • >95% of sequences mapped for human, mouse, and other key species • one of the largest collections RNA types of ncRNA genome annotations % of mapped human sequences Example of Rfam quality control and genome mapping Here is an Ensembl miRNA http://www.ensembl.org/Mus_musculus/Transcript/Summary?g=ENSMUSG00000106355;r=4:10874064-10874170;t =ENSMUST00000197675 But is it a miRNA or a tRNA? http://www.ensembl.org/Mus_musculus/Transcript/Summary?g=ENSMUSG00000106355;r=4:10874064-10874170;t =ENSMUST00000197675 RNAcentral shows a
  • Browsing Genes and Genomes with Ensembl

    Browsing Genes and Genomes with Ensembl

    Browsing Genes and Genomes with Ensembl Ensembl Workshop University of Cambridge 3rd September 2013 www.ensembl.org Denise Carvalho-Silva Ensembl Outreach Notes: This workshop is based on Ensembl release 72 (June 2013). 1) Presentation slides The pdf file of the talks presented in this workshop is available in the link below: http://www.ebi.ac.uk/~denise/naturimmun 2) Course booklet A diGital Copy of this course booklet can be found below: http://www.ebi.ac.uk/~denise/coursebooklet.pdf 3) Answers The answers to the exercises in this course booklet can be found below: http://www.ebi.ac.uk/~denise/answers 2 TABLE OF CONTENTS OVERVIEW ...................................................................................................... 4 INTRODUCTION TO ENSEMBL .................................................................. 5 BROWSER WALKTHROUGH .................................................................... 14 EXERCISES .................................................................................................... 38 BROWSER ..................................................................................................... 38 BIOMART ...................................................................................................... 41 VARIATION ................................................................................................... 48 COMPARATIVE GENOMICS ...................................................................... 50 REGULATION ..............................................................................................
  • Rfam 14: Expanded Coverage of Metagenomic, Viral and Microrna Families Ioanna Kalvari, Eric P

    Rfam 14: Expanded Coverage of Metagenomic, Viral and Microrna Families Ioanna Kalvari, Eric P

    Rfam 14: expanded coverage of metagenomic, viral and microRNA families Ioanna Kalvari, Eric P. Nawrocki, Nancy Ontiveros-Palacios, Joanna Argasinska, Kevin Lamkiewicz, Manja Marz, Sam Griffiths-Jones, Claire Toffano-Nioche, Daniel Gautheret, Zasha Weinberg, et al. To cite this version: Ioanna Kalvari, Eric P. Nawrocki, Nancy Ontiveros-Palacios, Joanna Argasinska, Kevin Lamkiewicz, et al.. Rfam 14: expanded coverage of metagenomic, viral and microRNA families. Nucleic Acids Research, 2020, 10.1093/nar/gkaa1047. hal-03031715 HAL Id: hal-03031715 https://hal.archives-ouvertes.fr/hal-03031715 Submitted on 14 Dec 2020 HAL is a multi-disciplinary open access L’archive ouverte pluridisciplinaire HAL, est archive for the deposit and dissemination of sci- destinée au dépôt et à la diffusion de documents entific research documents, whether they are pub- scientifiques de niveau recherche, publiés ou non, lished or not. The documents may come from émanant des établissements d’enseignement et de teaching and research institutions in France or recherche français ou étrangers, des laboratoires abroad, or from public or private research centers. publics ou privés. This article has been accepted for publication in Nucleic Acid Research Published by Oxford University Press : • DOI : 10.1093/nar/gkaa1047 • PUBMED : 33211869 Nucleic Acids Research, 2020 1 doi: 10.1093/nar/gkaa1047 Rfam 14: expanded coverage of metagenomic, viral and microRNA families Ioanna Kalvari 1,EricP.Nawrocki 2, Nancy Ontiveros-Palacios 1, Joanna Argasinska 1, Kevin Lamkiewicz 3,4, Manja Marz 3,4, Sam Griffiths-Jones 5, Claire Toffano-Nioche 6, Daniel Gautheret 6, Zasha Weinberg 7, Elena Rivas 8, Sean R. Eddy 8,9,10, Downloaded from https://academic.oup.com/nar/advance-article/doi/10.1093/nar/gkaa1047/5992291 by guest on 14 December 2020 Robert D.