Training Materials

Total Page:16

File Type:pdf, Size:1020Kb

Training Materials Training Materials • Ensembl training materials are protected by a CC BY license http://creativecommons.org/licenses/by/4.0/ • If you wish to re-use these materials, please credit Ensembl for their creation • If you use Ensembl for your work, please cite our papers http://www.ensembl.org/info/about/publications.html http://training.ensembl.org/events Browsing Genes and Genomes with Ensembl Astrid Gall Ensembl Outreach Officer, EMBL-EBI 13th Ensembl browser workshop Erasmus MC, Rotterdam EBI is an Outstation of the European Molecular Biology Laboratory. 13 - 14 February 2018 This Workshop Today Tomorrow Introduction & Exploring the Comparative genomics Ensembl genome browser Coffee/Tea Coffee/Tea Regulation Genes and transcripts Lunch break Lunch break BioMart Variation Coffee/Tea Coffee/Tea Tools, custom data, track The Variant Effect Predictor (VEP) hubs, advanced access http://training.ensembl.org/events Structure of the Workshop Presentation: What the data/tool is How we produce/process the data Demo: Follow along if Getting the data you want to Using the tool Exercises: Trying things out for yourself (alone/pairs?) Going beyond the demo Not a test! http://training.ensembl.org/events Extra Exercises Course Materials training.ensembl.org/events • Presentation • Coursebook (demos) • Questionbook (exercises) • Copysheet (text file for exercises) • Answerbook (exercise answers) http://training.ensembl.org/events Objectives You will learn: • What Ensembl is • What type of data you can explore in Ensembl • How to navigate the Ensembl browser website • Where to find help and documentation http://training.ensembl.org/events Questions? http://training.ensembl.org/events Introduction to Ensembl EBI is an Outstation of the European Molecular Biology Laboratory. Outline • Genome browsers • Ensembl features and access scales • Ensembl and Ensembl Genomes • Release Cycle • Genome Assembly http://training.ensembl.org/events Why do we need Genome Browsers? 1977: 1st genome sequenced (5 kb) 2004: Human genome sequence finished (3 Gb) http://training.ensembl.org/events Why do we need Genome Browsers? CGGCCTTTGGGCTCCGCCTTCAGCTCAAGACTTAACTTCCCTCCCAGCTGTCCCAGATGACGCCATCTGAAATTTCTTGGAAACACGATCAC TTTAACGGAATATTGCTGTTTTGGGGAAGTGTTTTACAGCTGCTGGGCACGCTGTATTTGCCTTACTTAAGCCCCTGGTAATTGCTGTATTC CGAAGACATGCTGATGGGAATTACCAGGCGGCGTTGGTCTCTAACTGGAGCCCTCTGTCCCCACTAGCCACGCGTCACTGGTTAGCGTGATT GAAACTAAATCGTATGAAAATCCTCTTCTCTAGTCGCACTAGCCACGTTTCGAGTGCTTAATGTGGCTAGTGGCACCGGTTTGGACAGCACA GCTGTAAAATGTTCCCATCCTCACAGTAAGCTGTTACCGTTCCAGGAGATGGGACTGAATTAGAATTCAAACAAATTTTCCAGCGCTTCTGA GTTTTACCTCAGTCACATAATAAGGAATGCATCCCTGTGTAAGTGCATTTTGGTCTTCTGTTTTGCAGACTTATTTACCAAGCATTGGAGGA ATATCGTAGGTAAAAATGCCTATTGGATCCAAAGAGAGGCCAACATTTTTTGAAATTTTTAAGACACGCTGCAACAAAGCAGGTATTGACAA ATTTTATATAACTTTATAAATTACACCGAGAAAGTGTTTTCTAAAAAATGCTTGCTAAAAACCCAGTACGTCACAGTGTTGCTTAGAACCAT AAACTGTTCCTTATGTGTGTATAAATCCAGTTAACAACATAATCATCGTTTGCAGGTTAACCACATGATAAATATAGAACGTCTAGTGGATA AAGAGGAAACTGGCCCCTTGACTAGCAGTAGGAACAATTACTAACAAATCAGAAGCATTAATGTTACTTTATGGCAGAAGTTGTCCAACTTT TTGGTTTCAGTACTCCTTATACTCTTAAAAATGATCTAGGACCCCCGGAGTGCTTTTGTTTATGTAGCTTACCATATTAGAAATTTAAAACT AAGAATTTAAGGCTGGGCGTGGTGGCTCACGCCTGTAATCCCAGCACTTTGGGAGGCCGAGGTGGGCGGATCACTTGAGGCCAGAAGTTTGA GACCAGCCTGGCCAACATGGTGAAACCCTATCTCTACTAAAAATACAAAAAATGTGCTGCGTGTGGTGGTGCGTGCCTGTAATCCCAGCTAC ACGGGAGGTGGAGGCAGGAGAATCGCTTGAACCCTGGAGGCAGAGGTTGCAGTGAGCCAAGATCATGCCACTGCACTCTAGCCTGGGCCACA TAGCATGACTCTGTCTCAAAACAAACAAACAAACAAAAAACTAAGAATTTAAAGTTAATTTACTTAAAAATAATGAAAGCTAACCCATTGCA TATTATCACAACATTCTTAGGAAAAATAACTTTTTGAAAACAAGTGAGTGGAATAGTTTTTACATTTTTGCAGTTCTCTTTAATGTCTGGCT AAATAGAGATAGCTGGATTCACTTATCTGTGTCTAATCTGTTATTTTGGTAGAAGTATGTGAAAAAAAATTAACCTCACGTTGAAAAAAGGA ATATTTTAATAGTTTTCAGTTACTTTTTGGTATTTTTCCTTGTACTTTGCATAGATTTTTCAAAGATCTAATAGATATACCATAGGTCTTTC CCATGTCGCAACATCATGCAGTGATTATTTGGAAGATAGTGGTGTTCTGAATTATACAAAGTTTCCAAATATTGATAAATTGCATTAAACTA TTTTAAAAATCTCATTCATTAATACCACCATGGATGTCAGAAAAGTCTTTTAAGATTGGGTAGAAATGAGCCACTGGAAATTCTAATTTTCA TTTGAAAGTTCACATTTTGTCATTGACAACAAACTGTTTTCCTTGCAGCAACAAGATCACTTCATTGATTTGTGAGAAAATGTCTACCAAAT TATTTAAGTTGAAATAACTTTGTCAGCTGTTCTTTCAAGTAAAAATGACTTTTCATTGAAAAAATTGCTTGTTCAGATCACAGCTCAACATG AGTGCTTTTCTAGGCAGTATTGTACTTCAGTATGCAGAAGTGCTTTATGTATGCTTCCTATTTTGTCAGAGATTATTAAAAGAAGTGCTAAA GCATTGAGCTTCGAAATTAATTTTTACTGCTTCATTAGGACATTCTTACATTAAACTGGCATTATTATTACTATTATTTTTAACAAGGACAC TCAGTGGTAAGGAATATAATGGCTACTAGTATTAGTTTGGTGCCACTGCCATAACTCATGCAAATGTGCCAGCAGTTTTACCCAGCATCATC TTTGCACTGTTGATACAAATGTCAACATCATGAAAAAGGGTTGAAAAAAGGAATATTTTAATAGTTTTCAGTTACTTTATGACTGTTAGCTA http://training.ensembl.org/events Genome Browsers unlock the Code https://www.ensembl.org https://www.ncbi.nlm.nih. gov/projects/mapview https://genome.ucsc.edu http://training.ensembl.org/events Ensembl Features • Genomes and gene builds for >100 species • Variation data • Alignments, gene trees, homologues • Regulatory build (ENCODE) • Gene expression and function • BioMart (data export) • Tools for data processing, e.g. VEP • Display your own data • Programmatic access via APIs • Completely Open Source (FTP, GitHub) http://training.ensembl.org/events Access Scales One by one Main browser Mobile site BioMart REST API Perl API VEP MySQL Groups FTP Whole genome http://training.ensembl.org/events Ensembl - 100+ Vertebrate and Model Organism Genomes http://training.ensembl.org/events Ensembl Genomes - Extending Ensembl across the Taxonomic Space www.ensembl.org www.ensemblgenomes.org ● Vertebrates ● Bacteria ● Fungi ● Metazoa ● Other representative species ● Plants ● Protists http://training.ensembl.org/events Ensembl and Ensembl Genomes Ensembl Ensembl Genomes Released 2000 2009 Species Vertebrates (fly, worm and Non-vertebrates (protists, yeast as outgroups) plants, fungi, metazoa, bacteria) Annotation By Ensembl In collaboration with the scientific communities URL www.ensembl.org www.ensemblgenomes.org http://training.ensembl.org/events Release Cycle New/updated interfaces 9091 AugDec 2017 Updated New regulation genome data assemblies ~3 months Updated variation Underlying data software updates Comparative Genomics including new genes Updated and genomes http://training.ensembl.org/events gene sets How does a Genome Assembly work? Sequence reads CGGCCTTTGGGCTCCGCCTTCAGCTCAAGA CAGCTGTCCCAGATGAC ACTTAACTTCCCTCCCAGCTGTCC GGGCTCCGCCTTCAGCTC TCCCAGCTGTCCCAGATGACGCCAT AACTTCCCTCCCAGCT CGGCCTTTGGGCTCC TCCGCCTTCAGCTCAAGACTTAACTTC CAGATGACGCC Match up overlaps CGGCCTTTGGGCTCCGCCTTCAGCTCAAGA AACTTCCCTCCCAGCT CAGATGACGCC TCCGCCTTCAGCTCAAGACTTAACTTC TCCCAGCTGTCCCAGATGACGCCAT GGGCTCCGCCTTCAGCTC ACTTAACTTCCCTCCCAGCTGTCC CGGCCTTTGGGCTCC CAGCTGTCCCAGATGAC Contig CGGCCTTTGGGCTCCGCCTTCAGCTCAAGACTTAACTTCCCTCCCAGCTGTCCCAGATGACGCCAT http://training.ensembl.org/events Cloning into Bacterial Artificial Chromosomes (BACs) BL Cut into fragments & clone into BACs BL001 BL002 BL003 BL004 Grow in bacterial cultures http://training.ensembl.org/events Genome Assembly: Tilepath, contigs, scaffolds, chromosome BACs from different BL001 IM004 individuals assembled, AL002 CM003 including overlaps: Tilepath Overlaps trimmed to give contigs. A run of contigs with no gaps is a scaffold. Genetic maps used to assemble scaffolds into a chromosome. http://training.ensembl.org/events Genome Contigs A contig is a contiguous stretch of DNA sequence that BL102 has been assembled based on sequencing information. BL AL476 AL CM553 CM IM768 IM http://training.ensembl.org/events Human Genome Assemblies • GRCh38 (aka hg38) • No gaps. Many rare/private alleles replaced. • www.ensembl.org • Most up-to-date and supported • GRCh37 (aka hg19) • 250 gaps • grch37.ensembl.org • Limited data and software updates • NCBI36 (aka hg18) • 150,000 gaps • ncbi36.ensembl.org • No longer updated http://training.ensembl.org/events Questions? http://training.ensembl.org/events Exploring the Ensembl Genome Browser: Homepage, Assemblies & Species EBI is an Outstation of the European Molecular Biology Laboratory. Hands on We will look at the Ensembl homepage and find information about the genome assemblies and species in Ensembl. • Demo: Coursebook page 7-13 • Exercises: Questionbook page 3 • Answers: Answerbook page 3-4 http://training.ensembl.org/events Questions? http://training.ensembl.org/events Exploring the Ensembl Genome Browser: The Location tab & Region in detail view EBI is an Outstation of the European Molecular Biology Laboratory. Hands on We will look at a region of the human genome, 4:122868000-122946000, and manipulate the view to see the data we are interested in. • Demo: Coursebook page 13-19 • Exercises: Questionbook page 3-5 • Answers: Answerbook page 4-8 http://training.ensembl.org/events Questions? http://training.ensembl.org/events Genes and Transcripts EBI is an Outstation of the European Molecular Biology Laboratory. Outline • Gene views in Ensembl • Annotation • Automatic • Manual • Imported • Golden transcripts, GENCODE, CCDS • Which transcript do I use? • Ensembl stable IDs • Gene Ontology http://training.ensembl.org/events Gene Views Coding exon Intron Non-coding exon Merged transcript Protein coding transcript Non-coding transcript http://training.ensembl.org/events Ensembl and Havana Annotation Automatic Annotation Manual Annotation http://training.ensembl.org/events Automatic Annotation • Genome-wide determination using the Ensembl automated pipeline • Predictions based on experimental (biological) data http://training.ensembl.org/events Biological Evidence International Nucleotide Sequence Database Collaboration • cDNAs • ESTs • RNAseq Protein sequence databases • Swiss-Prot: Manually curated • TrEMBL:
Recommended publications
  • Uniprot at EMBL-EBI's Role in CTTV
    Barbara P. Palka, Daniel Gonzalez, Edd Turner, Xavier Watkins, Maria J. Martin, Claire O’Donovan European Bioinformatics Institute (EMBL-EBI), European Molecular Biology Laboratory, Wellcome Genome Campus, Hinxton, Cambridge, CB10 1SD, UK UniProt at EMBL-EBI’s role in CTTV: contributing to improved disease knowledge Introduction The mission of UniProt is to provide the scientific community with a The Centre for Therapeutic Target Validation (CTTV) comprehensive, high quality and freely accessible resource of launched in Dec 2015 a new web platform for life- protein sequence and functional information. science researchers that helps them identify The UniProt Knowledgebase (UniProtKB) is the central hub for the collection of therapeutic targets for new and repurposed medicines. functional information on proteins, with accurate, consistent and rich CTTV is a public-private initiative to generate evidence on the annotation. As much annotation information as possible is added to each validity of therapeutic targets based on genome-scale experiments UniProtKB record and this includes widely accepted biological ontologies, and analysis. CTTV is working to create an R&D framework that classifications and cross-references, and clear indications of the quality of applies to a wide range of human diseases, and is committed to annotation in the form of evidence attribution of experimental and sharing its data openly with the scientific community. CTTV brings computational data. together expertise from four complementary institutions: GSK, Biogen, EMBL-EBI and Wellcome Trust Sanger Institute. UniProt’s disease expert curation Q5VWK5 (IL23R_HUMAN) This section provides information on the disease(s) associated with genetic variations in a given protein. The information is extracted from the scientific literature and diseases that are also described in the OMIM database are represented with a controlled vocabulary.
    [Show full text]
  • Ensembl Genomes: Extending Ensembl Across the Taxonomic Space P
    Published online 1 November 2009 Nucleic Acids Research, 2010, Vol. 38, Database issue D563–D569 doi:10.1093/nar/gkp871 Ensembl Genomes: Extending Ensembl across the taxonomic space P. J. Kersey*, D. Lawson, E. Birney, P. S. Derwent, M. Haimel, J. Herrero, S. Keenan, A. Kerhornou, G. Koscielny, A. Ka¨ ha¨ ri, R. J. Kinsella, E. Kulesha, U. Maheswari, K. Megy, M. Nuhn, G. Proctor, D. Staines, F. Valentin, A. J. Vilella and A. Yates EMBL-European Bioinformatics Institute, Wellcome Trust Genome Campus, Cambridge CB10 1SD, UK Received August 14, 2009; Revised September 28, 2009; Accepted September 29, 2009 ABSTRACT nucleotide archives; numerous other genomes exist in states of partial assembly and annotation; thousands of Ensembl Genomes (http://www.ensemblgenomes viral genomes sequences have also been generated. .org) is a new portal offering integrated access to Moreover, the increasing use of high-throughput genome-scale data from non-vertebrate species sequencing technologies is rapidly reducing the cost of of scientific interest, developed using the Ensembl genome sequencing, leading to an accelerating rate of genome annotation and visualisation platform. data production. This not only makes it likely that in Ensembl Genomes consists of five sub-portals (for the near future, the genomes of all species of scientific bacteria, protists, fungi, plants and invertebrate interest will be sequenced; but also the genomes of many metazoa) designed to complement the availability individuals, with the possibility of providing accurate and of vertebrate genomes in Ensembl. Many of the sophisticated annotation through the similarly low-cost databases supporting the portal have been built in application of functional assays.
    [Show full text]
  • Abstracts Genome 10K & Genome Science 29 Aug - 1 Sept 2017 Norwich Research Park, Norwich, Uk
    Genome 10K c ABSTRACTS GENOME 10K & GENOME SCIENCE 29 AUG - 1 SEPT 2017 NORWICH RESEARCH PARK, NORWICH, UK Genome 10K c 48 KEYNOTE SPEAKERS ............................................................................................................................... 1 Dr Adam Phillippy: Towards the gapless assembly of complete vertebrate genomes .................... 1 Prof Kathy Belov: Saving the Tasmanian devil from extinction ......................................................... 1 Prof Peter Holland: Homeobox genes and animal evolution: from duplication to divergence ........ 2 Dr Hilary Burton: Genomics in healthcare: the challenges of complexity .......................................... 2 INVITED SPEAKERS ................................................................................................................................. 3 Vertebrate Genomics ........................................................................................................................... 3 Alex Cagan: Comparative genomics of animal domestication .......................................................... 3 Plant Genomics .................................................................................................................................... 4 Ksenia Krasileva: Evolution of plant Immune receptors ..................................................................... 4 Andrea Harper: Using Associative Transcriptomics to predict tolerance to ash dieback disease in European ash trees ............................................................................................................
    [Show full text]
  • The ELIXIR Core Data Resources: ​Fundamental Infrastructure for The
    Supplementary Data: The ELIXIR Core Data Resources: fundamental infrastructure ​ for the life sciences The “Supporting Material” referred to within this Supplementary Data can be found in the Supporting.Material.CDR.infrastructure file, DOI: 10.5281/zenodo.2625247 (https://zenodo.org/record/2625247). ​ ​ Figure 1. Scale of the Core Data Resources Table S1. Data from which Figure 1 is derived: Year 2013 2014 2015 2016 2017 Data entries 765881651 997794559 1726529931 1853429002 2715599247 Monthly user/IP addresses 1700660 2109586 2413724 2502617 2867265 FTEs 270 292.65 295.65 289.7 311.2 Figure 1 includes data from the following Core Data Resources: ArrayExpress, BRENDA, CATH, ChEBI, ChEMBL, EGA, ENA, Ensembl, Ensembl Genomes, EuropePMC, HPA, IntAct /MINT , InterPro, PDBe, PRIDE, SILVA, STRING, UniProt ● Note that Ensembl’s compute infrastructure physically relocated in 2016, so “Users/IP address” data are not available for that year. In this case, the 2015 numbers were rolled forward to 2016. ● Note that STRING makes only minor releases in 2014 and 2016, in that the interactions are re-computed, but the number of “Data entries” remains unchanged. The major releases that change the number of “Data entries” happened in 2013 and 2015. So, for “Data entries” , the number for 2013 was rolled forward to 2014, and the number for 2015 was rolled forward to 2016. The ELIXIR Core Data Resources: fundamental infrastructure for the life sciences ​ 1 Figure 2: Usage of Core Data Resources in research The following steps were taken: 1. API calls were run on open access full text articles in Europe PMC to identify articles that ​ ​ mention Core Data Resource by name or include specific data record accession numbers.
    [Show full text]
  • PDF of Connecting Science's 2017
    Connecting Science Wellcome Genome Campus Hinxton, Cambridgeshire CB10 1RQ wellcomegenomecampus.org/connectingscience Annual Review 2017/18 2 02 Contents DIRECTO R’S INTRODUCTION Connecting Science in twelve months 04 – 05 Global ambitions Have you had your say on your DNA? 10 – 1 1 Increasing global impact by tailoring training needs 12 – 13 Worm hunting in Colombia 14 – 16 Catalysing collaboration Bringing together scientific partners 20 – 21 First global event for genetic counsellors 22 – 23 Decoding genomes together 24 – 26 Snapshot: May 2017 to April 2018 27 – 32 Democratising genomics CONNECTING SCIENCE Blood, sweat and success – Implementing a gender balance policy 36 – 37 2017/18 ANNUAL REVIEW Science communication: remembering to listen 38 – 39 Exploring our world in the Genome Gallery 40 – 41 The mission of Connecting Science is simple: to enable team, the commitment and support of our many partners and everyone to explore genomic science and its impact on collaborators, and the financial support and encouragement research, health and society. Behind those simple words is of Wellcome. I am personally very thankful for all of those Wellcome Genome Campus: considerable complexity. The science itself is complex and things, and know that this support and energy often translates ever-changing, with new technologies being continually into life- or career-changing experiences for the tens of A hub of knowledge, learning, developed and put into practice in both research and thousands of people we engage with directly every year. healthcare. The work that we do is also deceptively and engagement complex, reaching audiences from primary schools to I hope that some of the stories highlighted in this Annual research scientists, NHS staff to patients and their families; Review give you a sense of the excitement we have in all Creating spaces for thinking 46 – 47 it spans the globe, and everything we do involves constant that we do.
    [Show full text]
  • The Metagenomic Data Life-Cycle: Standards and Best Practices Petra Ten Hoopen1, Robert D
    GigaScience, 6, 2017, 1–11 doi: 10.1093/gigascience/gix047 Advance Access Publication Date: 16 June 2017 Review REVIEW Downloaded from https://academic.oup.com/gigascience/article-abstract/6/8/gix047/3869082 by guest on 01 October 2019 The metagenomic data life-cycle: standards and best practices Petra ten Hoopen1, Robert D. Finn1, Lars Ailo Bongo2, Erwan Corre3, Bruno Fosso4, Folker Meyer5, Alex Mitchell1, Eric Pelletier6,7,8, Graziano Pesole4,9, Monica Santamaria4, Nils Peder Willassen2 and Guy Cochrane1,∗ 1European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom, 2UiT The Arctic University of Norway, Tromsø N-9037, Norway, 3CNRS-UPMC, FR 2424, Station Biologique, Roscoff 29680, France, 4Institute of Biomembranes, Bioenergetics and Molecular Biotechnologies, CNR, Bari 70126, Italy, 5Argonne National Laboratory, Argonne IL 60439, USA, 6Genoscope, CEA, Evry´ 91000, France, 7CNRS/UMR-8030, Evry´ 91000, France, 8Universite´ Evry´ val d’Essonne, Evry´ 91000, France and 9Department of Biosciences, Biotechnologies and Biopharmaceutics, University of Bari “A. Moro,” Bari 70126, Italy ∗Correspondence address. Guy Cochrane, European Molecular Biology Laboratory, European Bioinformatics Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, UK. Tel: +44(0)1223-494444; E-mail: [email protected] Abstract Metagenomics data analyses from independent studies can only be compared if the analysis workflows are described ina harmonized way. In this overview, we have mapped the landscape of data standards available for the description of essential steps in metagenomics: (i) material sampling, (ii) material sequencing, (iii) data analysis, and (iv) data archiving and publishing. Taking examples from marine research, we summarize essential variables used to describe material sampling processes and sequencing procedures in a metagenomics experiment.
    [Show full text]
  • Annual Scientific Report 2013 on the Cover Structure 3Fof in the Protein Data Bank, Determined by Laponogov, I
    EMBL-European Bioinformatics Institute Annual Scientific Report 2013 On the cover Structure 3fof in the Protein Data Bank, determined by Laponogov, I. et al. (2009) Structural insight into the quinolone-DNA cleavage complex of type IIA topoisomerases. Nature Structural & Molecular Biology 16, 667-669. © 2014 European Molecular Biology Laboratory This publication was produced by the External Relations team at the European Bioinformatics Institute (EMBL-EBI) A digital version of the brochure can be found at www.ebi.ac.uk/about/brochures For more information about EMBL-EBI please contact: [email protected] Contents Introduction & overview 3 Services 8 Genes, genomes and variation 8 Molecular atlas 12 Proteins and protein families 14 Molecular and cellular structures 18 Chemical biology 20 Molecular systems 22 Cross-domain tools and resources 24 Research 26 Support 32 ELIXIR 36 Facts and figures 38 Funding & resource allocation 38 Growth of core resources 40 Collaborations 42 Our staff in 2013 44 Scientific advisory committees 46 Major database collaborations 50 Publications 52 Organisation of EMBL-EBI leadership 61 2013 EMBL-EBI Annual Scientific Report 1 Foreword Welcome to EMBL-EBI’s 2013 Annual Scientific Report. Here we look back on our major achievements during the year, reflecting on the delivery of our world-class services, research, training, industry collaboration and European coordination of life-science data. The past year has been one full of exciting changes, both scientifically and organisationally. We unveiled a new website that helps users explore our resources more seamlessly, saw the publication of ground-breaking work in data storage and synthetic biology, joined the global alliance for global health, built important new relationships with our partners in industry and celebrated the launch of ELIXIR.
    [Show full text]
  • Strategic Plan 2011-2016
    Strategic Plan 2011-2016 Wellcome Trust Sanger Institute Strategic Plan 2011-2016 Mission The Wellcome Trust Sanger Institute uses genome sequences to advance understanding of the biology of humans and pathogens in order to improve human health. -i- Wellcome Trust Sanger Institute Strategic Plan 2011-2016 - ii - Wellcome Trust Sanger Institute Strategic Plan 2011-2016 CONTENTS Foreword ....................................................................................................................................1 Overview .....................................................................................................................................2 1. History and philosophy ............................................................................................................ 5 2. Organisation of the science ..................................................................................................... 5 3. Developments in the scientific portfolio ................................................................................... 7 4. Summary of the Scientific Programmes 2011 – 2016 .............................................................. 8 4.1 Cancer Genetics and Genomics ................................................................................ 8 4.2 Human Genetics ...................................................................................................... 10 4.3 Pathogen Variation .................................................................................................. 13 4.4 Malaria
    [Show full text]
  • Browsing Genomes with Ensembl Annotation
    Browsing genomes with EnsEMBL Annotation • During recent years release of large amounts of sequence data • Raw sequence data are not so useful on its own. They are most valuable when provided with comprehensive good quality annotation CCCAACAAGAATGTAAAATCTTTAAGTGCCTGTTTTCATACTTATTTGACCACCCTATCTCTAGAATCTTGCATGATG TCTAGCCCTAGTAGGATCAAAAAATACTTACAAAGCAACTGAATAGCTACATGAATAGATGGATGAATAAATGCATG GGTGGATGGATGGATTAATGAAATCATTTATATGACTTAAAGTTTGCAGAGGAGTATCATATTTGGAAGGCAGTAAG GAAGTCTGTGTAGTCGATGGTAAAGGCAATTGGGAAGTTTGTTAGGCACAATAGGTCAAAATTTGTTTTTGAAGTCC TGTTACTTCACGTTTCTTTGTTTCACTTTCTTAAAACAGGAAACTCTTTTCTATGATCATTCTTCCAGGGCCTGGCTCT TCATCTGCAACCCAGTAATATCCCTAATGTCAAAAAGCTACTGGTTTAATTCGTGCCATTTTCAAAGAGGACTACTGA ATTCTGATGTGGCTTCAAACATTTAGGTTAGGCATATCTAATGGAGAACTTGCAGCCACACTGACTTGTAGTGAAAT ATCTATTTTGAGCCTGCCCAGTGTTGCTTAAATTGTAGTTTTCCTTGCCAGCTATTCATACAAGAGATGTGAGAAGCA CCATAAAAGGCGTTGTGAGGAGTTGTGGGGGAGTGAGGGAGAGAAGAGGTTGAAAAGCTTATTAGCTGCTGTACGG TAAAAGTGAGCTCTTACGGGAATGGGAATGTAGTTTTAGCCCTCCAGGGATTCTATTTAGCCCGCCAGGAATTAACC TTGACTATAAATAGGCCATCAATGACCTTTCCAGAGAATGTTCAGAGACCTCAACTTTGTTTAGAGATCTTGTGTGGG TGGAACTTCCTGTTTGCACACAGAGCAGCATAAAGCCCAGTTGCTTTGGGAAGTGTTTGGGACCAGATGGATTGTAG GGAGTAGGGTACAATACAGTCTGTTCTCCTCCAGCTCCTTCTTTCTGCAACATGGGGAAGAACAAACTCCTTCATCC AAGTCTGGTTCTTCTCCTCTTGGTCCTCCTGCCCACAGACGCCTCAGTCTCTGGAAAACCGTGAGTTCCACACAGAG AGCGTGAAGCATGAACCTAGAGTCCTTCATTTATTGCAGATTTTTCTTTATATCATTCCTTTTTCTTTCCTATGATACT GTCATCTTCTTATCTCTAAGATTCCTTCCAGATTTTACAAATCTAGTTTACTCATTACTTGCTTACTTTTAATCATTCT TCCCCAACTCTCTGAAGCTCTAATATGCAAAGCCTTCCTAAGGGGTGTCAGAAATTTTTAGCTTTTTAAAAGAATAAA
    [Show full text]
  • The Genomic Basis of Circadian and Circalunar Timing Adaptations in a Midge Tobias S
    OPEN ARTICLE doi:10.1038/nature20151 The genomic basis of circadian and circalunar timing adaptations in a midge Tobias S. Kaiser1,2,3†, Birgit Poehn1,3, David Szkiba2, Marco Preussner4, Fritz J. Sedlazeck2†, Alexander Zrim2, Tobias Neumann1,2, Lam-Tung Nguyen2,5, Andrea J. Betancourt6, Thomas Hummel3,7, Heiko Vogel8, Silke Dorner1, Florian Heyd4, Arndt von Haeseler2,3,5 & Kristin Tessmar-Raible1,3 Organisms use endogenous clocks to anticipate regular environmental cycles, such as days and tides. Natural variants resulting in differently timed behaviour or physiology, known as chronotypes in humans, have not been well characterized at the molecular level. We sequenced the genome of Clunio marinus, a marine midge whose reproduction is timed by circadian and circalunar clocks. Midges from different locations show strain-specific genetic timing adaptations. We examined genetic variation in five C. marinus strains from different locations and mapped quantitative trait loci for circalunar and circadian chronotypes. The region most strongly associated with circadian chronotypes generates strain-specific differences in the abundance of calcium/calmodulin-dependent kinase II.1 (CaMKII.1) splice variants. As equivalent variants were shown to alter CaMKII activity in Drosophila melanogaster, and C. marinus (Cma)-CaMKII.1 increases the transcriptional activity of the dimer of the circadian proteins Cma-CLOCK and Cma-CYCLE, we suggest that modulation of alternative splicing is a mechanism for natural adaptation in circadian timing. Around the new or full moon, during a few specific hours surround- Our study aimed to identify the genetic basis of C. marinus adaptation ing low tide, millions of non-biting midges of the species C.
    [Show full text]
  • ALEXA: a Microarray Design Platform for Alternative Expression Analysis
    CORRESPONDENCE expression of alternative mRNA isoforms in 5-fluorouracil (5-FU)- ALEXA: a microarray design platform for sensitive and resistant colorectal cancer cell lines5 and compared alternative expression analysis the results to those from the Affymetrix ‘GeneChip Human Exon 1.0 ST’ array (see Supplementary Results, Supplementary Fig. 2 To the editor: Eukaryotic genomes are predicted to contain about and Supplementary Table 2 online). Genes and exons differentially 7,000–29,000 genes1. Each of these genes may be alternatively expressed between 5-FU–sensitive and resistant cells were identi- processed to produce multiple distinct mRNAs by alternative fied by both platforms (with significant overlap), but ALEXA arrays transcript initiation, splicing and polyadenylation (collectively provided additional information on the connectivity and boundar- referred to as alternative expression). Although analysis of avail- ies of exons (Table 1). Furthermore, alternative expression events able transcript resources indicates that up to ~75% of genes are identified by ALEXA were significantly enriched for known alterna- alternatively processed, most microarray expression platforms tive expression events represented in publicly available mRNA and cannot detect alternative transcripts2. expressed sequence tag (EST) databases (Supplementary Results methods Proof-of-principle experiments have described the use of oli- and Supplementary Data 1 online). Finally, we demonstrated the gonucleotide microarrays to profile transcript isoforms gener- advantage of the ALEXA approach by identifying several differen- ated by alternative expression, but resources to create such arrays tially expressed known and predicted isoforms with potential rele- are lacking3,4. To address this limitation we created a microarray vance to 5-FU resistance (Supplementary Fig. 3 and Supplementary .com/nature e design platform for alternative expression analysis (ALEXA), Tables 3 and 4 online).
    [Show full text]
  • 5 February 2020 Draft Programme
    Genomic Practice for Genetic Counsellors Wellcome Genome Campus Hinxton, Cambridge, UK 3 - 5 February 2020 Draft Programme Monday 3 February 11:00 – 11:00 Registration with coffee 11:30 – 12:00 Welcome and introduction to the course 12:00 – 13:45 Session 1: The role of genomics in healthcare The role in the NHS Nicki Taverner Cardiff University and All Wales Medical Genetic Service, UK Catherine Houghton Liverpool Women's NHS Foundation Trust, UK Outside the NHS Gemma Chandratilake University of Cambridge, UK An international perspective – Melbourne Genomics Health Alliance Lyndon Gallacher Victorian Clinicial Genetics Service, Australia Q&A with speakers 13:45 – 14:45 Lunch 14:45 – 16:00 Session 2: Variant interpretation Introduction to a genome browser Gemma Chandratillake University of Cambridge, UK Variant interpretation: going from millions to one of interest that could be the answer Helen Firth Cambridge University Hospitals, UK 16:00 – 16:20 Afternoon tea 16:20 – 18.20 Workshop: Variant interpretation using DECIPHER {and other approaches} Julia Foreman WSI, UK + Gemma Chandratillake University of Cambridge, UK 18:29 – 19:00 Consolidating learning for Day 1, Q+A Led by programme committee 19:00 Dinner Tuesday 4 February 09:00 – 10:00 Session 3: Cancer genomics Cancer Genomics: bridging from the tumour to the germline in variant interpretation Clare Turnbull The Institute of Cancer Research, UK 10:00 – 12:00 Workshop: Cancer variant interpretation Heather Pierce Cambridge University Hospitals, UK 12:00 – 13:00 Functional studies
    [Show full text]