Analysis of Copy Loss and Gain Variations in Holstein Cattle

Total Page:16

File Type:pdf, Size:1020Kb

Analysis of Copy Loss and Gain Variations in Holstein Cattle Seroussi et al. BMC Genomics 2010, 11:673 http://www.biomedcentral.com/1471-2164/11/673 RESEARCH ARTICLE Open Access Analysis of copy loss and gain variations in Holstein cattle autosomes using BeadChip SNPs Eyal Seroussi1*, Giora Glick1, Andrey Shirak1, Emanuel Yakobson1, Joel I Weller1, Ephraim Ezra2, Yoel Zeron3 Abstract Background: Copy number variation (CNV) has been recently identified in human and other mammalian genomes, and there is a growing awareness of CNV’s potential as a major source for heritable variation in complex traits. Genomic selection is a newly developed tool based on the estimation of breeding values for quantitative traits through the use of genome-wide genotyping of SNPs. Over 30,000 Holstein bulls have been genotyped with the Illumina BovineSNP50 BeadChip, which includes 54,001 SNPs (~SNP/50,000 bp), some of which fall within CNV regions. Results: We used the BeadChip data obtained for 912 Israeli bulls to investigate the effects of CNV on SNP calls. For each of the SNPs, we estimated the frequencies of occurrence of loss of heterozygosity (LOH) and of gain, based either on deviation from the expected Hardy-Weinberg equilibrium (HWE) or on signal intensity (SI) using the PennCNV “detect” option. Correlations between LOH/CNV frequencies predicted by the two methods were low (up to r = 0.08). Nevertheless, 418 locations displayed significantly high frequencies by both methods. Efficiency of designating large genomic clusters of olfactory receptors as CNVs was 29%. Frequency values for copy loss were distinguishable in non-autosomal regions, indicating misplacement of a region in the current BTA7 map. Analysis of BTA18 placed major quantitative trait loci affecting net merit in the US Holstein population in regions rich in segmental duplications and CNVs. Enrichment of transporters in CNV loci suggested their potential effect on milk- production traits. Conclusions: Expansion of HWE and PennCNV analyses allowed estimating LOH/CNV frequencies, and combining the two methods yielded more sensitive detection of inherited CNVs and better estimation of their possible effects on cattle genetics. Although this approach was more effective than methodologies previously applied in cattle, it has severe limitations. Thus the number of CNVs reported here for the Holstein breed may represent as little as one-tenth of inherited common structural variation. Background (~SNP/50,000 bp). This chip may capture any genetic The Holstein-Friesian breed is the world’shighest- variance that is genetically linked to these markers, as producing dairy cattle; much of its outstanding milk well as copy number variations (CNVs) [2,3]. A CNV is production was gained by selection of elite artificial a structural variation, including deletion, duplication, insemination (AI) bulls based on breeding values that translocation or inversion. CNV has been recently iden- were estimated by progeny testing. Genomic selection is tified in human and other mammalian genomes, and it a newly developed tool for the estimation of breeding is now recognized that CNV might be a major source of values through the use of genome-wide genotyping of heritable variation in complex traits [4]. In humans, single nucleotide polymorphisms (SNPs). Over 30,000 over 14,478 CNV loci have been recorded based on Holstein bulls have been genotyped with the Illumina 89,427 different entries that cover about one-third of BovineSNP50 BeadChip [1], which includes 54,001 SNPs the genome. Of these entries, 65% include CNVs that range mostly between 1 and 10 kb and 34% are indels * Correspondence: [email protected] in the range of 100 bp to 1 kb http://projects.tcag.ca/ 1Institute of Animal Sciences, ARO, The Volcani Center, Bet Dagan 50250, variation/. CNV regions (CNVRs) encompassing adja- Israel cent or overlapping losses or gains cover 12% of the Full list of author information is available at the end of the article © 2010 Seroussi et al; licensee BioMed Central Ltd. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/2.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. Seroussi et al. BMC Genomics 2010, 11:673 Page 2 of 10 http://www.biomedcentral.com/1471-2164/11/673 human genome. Hence, this source of variation has divided by n, the solution for this expansion in the case 0.5 more nucleotide content per genome than SNPs [4]. of a null allele r is: rl = [0.25 - 0.25pqo +poqo/pqo] - However, assuming an average spontaneous CNV muta- 0.5; in the case of gain of an allele which consists of 0.5 0.5 0.5 tion rate of 1/10,000 per locus [5], it is expected that a both types, r is: rg =[po +qo +pqo] -po -qo (see considerable portion of the reported entries arise from Additional file 1 for a detailed mathematical solution de novo CNVs of a sporadic nature. and Additional file 2 for allele distribution, c2 test, and Several algorithms for CNV identification from SNP rl and rg values for all 54,001 SNPs of the Illumina Bovi- arrays are available [6]. Following reports that PennCNV neSNP50 BeadChip). Average values of rl and rg for was the most reliable algorithm in the detection of autosomal markers were -0.3% ± 2.7 and 0.5% ± 2.1, CNVs from Illumina BeadChip data [7,8], we chose this respectively. For non-pseudoautosomal markers on the software to analyze signal intensity (SI) data. PennCNV X chromosome (positions 28,044-86,115,497), where the is a CNV detection tool that incorporates multiple model of inheritance does not fit the model under sources of information, including the ratio of total SI to which the formulas were developed, average values of rl allelic intensity at each SNP marker. This software was and rg were 361% ± 258 and -25% ± 16, respectively. originally developed for Illumina whole-genome Bead- Thus, encountering extreme values (above 100%) for rl, Chip arrays [9]. when analyzing autosomal markers, may indicate an The introduction of AI to modern dairy herd manage- error in the mapping of markers that are actually ment has resulted in a loss of genetic diversity in Hol- located on sex chromosomes. Although the HWE devia- steins and the effective size of the Holstein population is tion is an important factor in CNV occurrence, other low (e.g. 39 in the USA [10]). Since the Israeli Holstein reasons than erroneous positioning of markers of sex population has been under intensive selection for chromosomes may also exist; for example, systematic 50 years, its genetic pool is expected to have similar problems in distinguishing the alleles, due to technical characteristics. Although it is now accepted that gen- failures. However these are unlikely, because of the high omes vary more at the structural level than at the quality of the Beadchip technology [11]. nucleotide-sequence level, little has been published on CNVs in Holsteins. SI-based detection of CNV In a study that validated the quality of BovineSNP50 Using the PennCNV detection module, we analyzed the BeadChip performance [11], population-wide genotyping autosomes of each of the 912 bulls for CNVs. From the of Israeli Holstein bulls was initiated in order to intro- output of this analysis, which contained the chromoso- duce genomic selection into the Israeli breeding program. mal positions and copy numbers of the detected CNVs, Our study makes use of these data to describe the the frequency of loss or gain for each SNP marker was frequent CNV in Holsteins and investigate its effect on calculated (Additional file 2). Average loss and gain BeadChip calls. We propose to combine Hardy-Weinberg values for autosomal markers were lsi = 0.02% ± 0.2 and equilibrium (HWE)-based and SI-based methods to reli- gsi = 0.12% ± 0.15, respectively. ably detect CNVs of the deletion and duplication types that are not de novo or sporadic CNVs, which are less Comparing the HWE-based and SI-based methods for likely to be of any economic value. CNV detection Seeking confirmation of the CNV detection, we examined Results and Discussion the correlation between the HWE-based and SI-based HWE-based detection of CNV detection methods (Figure 1). When sorted according to We used the data obtained for 912 Holstein bulls to frequency of deletion as detected by the PennCNV analy- investigate the effects of CNVs on BovineSNP50 Bead- sis (SI loss), the markers that exhibited frequent LOH Chip calls. For each of the SNPs, we estimated the fre- using the HWE-based formulas (high rl, and negative rg) quency of occurrence of deletions and insertions using a clustered together towards the right end of panel A generalization of the Hardy-Weinberg principle for (Figure 1). The distribution of frequencies of the SI-based more than two allele frequencies (p, q) by assuming pre- method was of limited range (from 0 to 14%) compared sence of a third allele (r). Under the assumption of to the HWE-based method (Figure 1A). Six autosomal three-alleles, expected HWE frequencies are obtained by markers displayed rl values higher than 100%, five of the trinomial expansion of (p + q+ r)2 = 1. Defining ‘n’ them that were closely mapped on BTA7 (see BTA7 sec- as the number of individuals sampled, ‘pqo’ as the tion). These markers were regarded as non-autosomal detected number of individuals with phenotype similar and removed from further analyses. to the pq heterozygote phenotype divided by n, ‘po’ as As the Holstein population has a very low effective thenumberofallelep-likehomozygotesdividedbyn, population size, it was expected that some of the CNV and ‘qo’ as the number of allele q-like homozygotes alleles within our sample would be common.
Recommended publications
  • Essential Genes Shape Cancer Genomes Through Linear Limitation of Homozygous Deletions
    ARTICLE https://doi.org/10.1038/s42003-019-0517-0 OPEN Essential genes shape cancer genomes through linear limitation of homozygous deletions Maroulio Pertesi1,3, Ludvig Ekdahl1,3, Angelica Palm1, Ellinor Johnsson1, Linnea Järvstråt1, Anna-Karin Wihlborg1 & Björn Nilsson1,2 1234567890():,; The landscape of somatic acquired deletions in cancer cells is shaped by positive and negative selection. Recurrent deletions typically target tumor suppressor, leading to positive selection. Simultaneously, loss of a nearby essential gene can lead to negative selection, and introduce latent vulnerabilities specific to cancer cells. Here we show that, under basic assumptions on positive and negative selection, deletion limitation gives rise to a statistical pattern where the frequency of homozygous deletions decreases approximately linearly between the deletion target gene and the nearest essential genes. Using DNA copy number data from 9,744 human cancer specimens, we demonstrate that linear deletion limitation exists and exposes deletion-limiting genes for seven known deletion targets (CDKN2A, RB1, PTEN, MAP2K4, NF1, SMAD4, and LINC00290). Downstream analysis of pooled CRISPR/Cas9 data provide further evidence of essentiality. Our results provide further insight into how the deletion landscape is shaped and identify potentially targetable vulnerabilities. 1 Hematology and Transfusion Medicine Department of Laboratory Medicine, BMC, SE-221 84 Lund, Sweden. 2 Broad Institute, 415 Main Street, Cambridge, MA 02142, USA. 3These authors contributed equally: Maroulio Pertesi, Ludvig Ekdahl. Correspondence and requests for materials should be addressed to B.N. (email: [email protected]) COMMUNICATIONS BIOLOGY | (2019) 2:262 | https://doi.org/10.1038/s42003-019-0517-0 | www.nature.com/commsbio 1 ARTICLE COMMUNICATIONS BIOLOGY | https://doi.org/10.1038/s42003-019-0517-0 eletion of chromosomal material is a common feature of we developed a pattern-based method to identify essential genes Dcancer genomes1.
    [Show full text]
  • Patient DNA Cross-Reactivity of the CDC SARS-Cov-2 Extraction Control Leads to an Inherent Potential for False Negative Results Authors: Adam P
    bioRxiv preprint doi: https://doi.org/10.1101/2020.05.13.094839; this version posted May 15, 2020. The copyright holder for this preprint (which was not certified by peer review) is the author/funder. All rights reserved. No reuse allowed without permission. Title: Patient DNA cross-reactivity of the CDC SARS-CoV-2 extraction control leads to an inherent potential for false negative results Authors: Adam P. Rosebrock1,2* Affiliations: 1Department of Pathology, Stony Brook University Hospital, Stony Brook, NY, USA. 2Stony Brook University Cancer Center, Stony Brook, NY, USA. *Correspondence to: [email protected] Abstract: Testing for RNA viruses such as SARS-CoV-2 requires careful handling of inherently labile RNA during sample collection, clinical processing, and molecular analysis. Tests must include fail-safe controls that affirmatively report the presence of intact RNA and demonstrate success of all steps of the assay. A result of “no virus signal” is insufficient for clinical interpretation: controls must also say “The reaction worked as intended and would have found virus if present.” Unfortunately, a widely used test specified by the US Centers for Disease Control and Prevention (CDC) incorporates a control that does not perform as intended and claimed. Detecting SARS-CoV-2 with this assay requires both intact RNA and successful reverse transcription. The CDC-specified control does not require either of these, due to its inability to differentiate human genomic DNA from reverse-transcribed RNA. Patient DNA is copurified from nasopharyngeal swabs during clinically-approved RNA extraction and is sufficient to return an “extraction control success” signal using the CDC design.
    [Show full text]
  • Systematic Discovery of Endogenous Human Ribonucleoprotein Complexes
    bioRxiv preprint doi: https://doi.org/10.1101/480061; this version posted November 27, 2018. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under aCC-BY 4.0 International license. Systematic discovery of endogenous human ribonucleoprotein complexes Anna L. Mallam1,2,3,§,*, Wisath Sae-Lee1,2,3, Jeffrey M. Schaub1,2,3, Fan Tu1,2,3, Anna Battenhouse1,2,3, Yu Jin Jang1, Jonghwan Kim1, Ilya J. Finkelstein1,2,3, Edward M. Marcotte1,2,3,§, Kevin Drew1,2,3,§,* 1 Department of Molecular Biosciences, University of Texas at Austin, Austin, TX 78712, USA 2 Center for Systems and Synthetic Biology, University of Texas at Austin, Austin, TX 78712, USA 3 Institute for Cellular and Molecular Biology, University of Texas at Austin, Austin, TX 78712, USA § Correspondence: [email protected] (A.L.M.), [email protected] (E.M.M.), [email protected] (K.D.) * These authors contributed equally to this work Short title: A resource of human ribonucleoprotein complexes Summary: An exploration of human protein complexes in the presence and absence of RNA reveals endogenous ribonucleoprotein complexes ! 1! bioRxiv preprint doi: https://doi.org/10.1101/480061; this version posted November 27, 2018. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under aCC-BY 4.0 International license. Abstract Ribonucleoprotein (RNP) complexes are important for many cellular functions but their prevalence has not been systematically investigated.
    [Show full text]
  • Droplet Digital™ PCR Applications Guide Table of Contents
    Droplet Digital PCR Droplet Digital™ PCR Applications Guide Table of Contents Chapter 1 Droplet Digital™ PCR . 1 Introduction . 1 QX100/QX200 Workflow . .2 Droplet Generation . 3 PCR Amplification . .4 Droplet Reading . 4 ddPCR Data Analysis . 5 Emerging Applications of Droplet Digital PCR . .7 ddPCR for Absolute Quantification and Experimental Considerations . 7 Chapter 2 Designing Droplet Digital™ PCR Experiments . 11 Assay Design for Droplet Digital PCR . 11 Designing Primers . .11 Designing Probes . 12 Designing an Assay . .13 Sample Preparation . 15 Adding DNA to the Reaction Mix . .16 ddPCR Experimental Workflow . .17 Droplet Generation . 17 PCR . 18 Setting Up an Experiment in QuantaSoft™ Software . .19 Droplet Reading . 20 Data Analysis . .20 Merging Wells . 22 PCR Optimization Using Thermal Gradients . 23 ddPCR Using the QX200 System and EvaGreen dsDNA Dye . .24 EvaGreen and Gene Expression . 26 Multiplexing with EvaGreen . 27 Reference . .27 Chapter 3 Absolute Quantification and the Statistics of Droplet Digital™ PCR . 28 Running Absolute Quantification Experiments . .28 Absolute Quantification Data Analysis . 29 iii | Droplet Digital™ PCR Applications Guide Table of Contents Statistics of ddPCR . .30 Copies per Microliter . .31 Copies per Droplet . 31 Low Concentration Example . .32 Intermediate Concentration Example . 33 High Concentration Example . 33 Looking across the Whole Concentration Range . .34 Concentration Calculation . .34 Definitions . .34 Formula for Calculating Concentration . .35 Derivation of Concentration Formula . 35 Errors in ddPCR . 36 Chapter 4 Copy Number Variation Analysis . .37 Overview . .37 CNV Calculations . 39 CNV Analysis in Homogeneous Samples . 39 CNV Analysis in Heterogeneous Samples . .39 Planning CNV Experiments . .40 Assay Design . 40 Running a CNV Assay . .41 Restriction Digestion . 41 DNA Loading for Lower-Order CN Analysis (diploid CN <10) .
    [Show full text]
  • The Neurodegenerative Diseases ALS and SMA Are Linked at The
    Nucleic Acids Research, 2019 1 doi: 10.1093/nar/gky1093 The neurodegenerative diseases ALS and SMA are linked at the molecular level via the ASC-1 complex Downloaded from https://academic.oup.com/nar/advance-article-abstract/doi/10.1093/nar/gky1093/5162471 by [email protected] on 06 November 2018 Binkai Chi, Jeremy D. O’Connell, Alexander D. Iocolano, Jordan A. Coady, Yong Yu, Jaya Gangopadhyay, Steven P. Gygi and Robin Reed* Department of Cell Biology, Harvard Medical School, 240 Longwood Ave. Boston MA 02115, USA Received July 17, 2018; Revised October 16, 2018; Editorial Decision October 18, 2018; Accepted October 19, 2018 ABSTRACT Fused in Sarcoma (FUS) and TAR DNA Binding Protein (TARDBP) (9–13). FUS is one of the three members of Understanding the molecular pathways disrupted in the structurally related FET (FUS, EWSR1 and TAF15) motor neuron diseases is urgently needed. Here, we family of RNA/DNA binding proteins (14). In addition to employed CRISPR knockout (KO) to investigate the the RNA/DNA binding domains, the FET proteins also functions of four ALS-causative RNA/DNA binding contain low-complexity domains, and these domains are proteins (FUS, EWSR1, TAF15 and MATR3) within the thought to be involved in ALS pathogenesis (5,15). In light RNAP II/U1 snRNP machinery. We found that each of of the discovery that mutations in FUS are ALS-causative, these structurally related proteins has distinct roles several groups carried out studies to determine whether the with FUS KO resulting in loss of U1 snRNP and the other two members of the FET family, TATA-Box Bind- SMN complex, EWSR1 KO causing dissociation of ing Protein Associated Factor 15 (TAF15) and EWS RNA the tRNA ligase complex, and TAF15 KO resulting in Binding Protein 1 (EWSR1), have a role in ALS.
    [Show full text]
  • A Quantitative Approach for Measuring the Reservoir of Latent HIV-1 Proviruses Katherine M
    LETTER https://doi.org/10.1038/s41586-019-0898-8 A quantitative approach for measuring the reservoir of latent HIV-1 proviruses Katherine M. Bruner1,8,10, Zheng Wang1,10, Francesco R. Simonetti1, Alexandra M. Bender1, Kyungyoon J. Kwon1, Srona Sengupta1, Emily J. Fray1, Subul A. Beg1, Annukka A. R. Antar1, Katharine M. Jenike1, Lynn N. Bertagnolli1, Adam A. Capoferri1, Joshua T. Kufera1, Andrew Timmons1, Christopher Nobles2, John Gregg2, Nikolas Wada3, Ya-Chi Ho1,9, Hao Zhang4, Joseph B. Margolick4, Joel N. Blankson1, Steven G. Deeks5, Frederic D. Bushman2, Janet D. Siliciano1, Gregory M. Laird6 & Robert F. Siliciano1,7* A stable latent reservoir for HIV-1 in resting CD4+ T cells is the efficacy of a cure intervention causing a selective one log reduction in principal barrier to a cure1–3. Curative strategies that target the intact proviruses may be inapparent (Fig. 1h). reservoir are being tested4,5 and require accurate, scalable reservoir Interrogating individual proviruses simultaneously at several posi- assays. The reservoir was defined with quantitative viral outgrowth tions could differentiate intact from defective proviruses. Analysis of assays for cells that release infectious virus after one round of T nFGS data revealed that strategically placed amplicons in the packaging cell activation1. However, these quantitative outgrowth assays and signal (Ψ) and env regions could jointly identify more than 90% of newer assays for cells that produce viral RNA after activation6 may deleted proviruses as defective (Fig. 2a). Hypermutated proviruses must underestimate the reservoir size because one round of activation also be identified (Fig. 1a). Seventy-three per cent of these have muta- does not induce all proviruses7.
    [Show full text]
  • Detecting Long Non-Coding RNA Biomarkers in Prostate Cancer Liquid Biopsies: Hype Or Hope?
    CORE Metadata, citation and similar papers at core.ac.uk Provided by Ghent University Academic Bibliography Non-coding RNA Research 3 (2018) 64e74 Contents lists available at ScienceDirect Non-coding RNA Research journal homepage: http://www.keaipublishing.com/NCRNA Detecting long non-coding RNA biomarkers in prostate cancer liquid biopsies: Hype or hope? * Hetty Helsmoortel a, b, , Celine Everaert a, b, Nicolaas Lumen b, c, Piet Ost b, d, Jo Vandesompele a, b a Center for Medical Genetics, Ghent University, Ghent, Belgium b Cancer Research Institute Ghent (CRIG), Ghent University, Ghent, Belgium c Department of Urology, Ghent University Hospital, Ghent, Belgium d Department of Radiation Oncology and Experimental Cancer Research, Ghent University, Ghent, Belgium article info abstract Article history: Prostate cancer is a heterogeneous malignancy, with clinical courses widely differing between indolent Received 3 January 2018 and aggressive lethal disease. This heterogeneity calls for a more personalized approach towards diag- Received in revised form nosis, prognosis, treatment decision, monitoring and follow-up of patients. In this review, we discuss the 2 May 2018 possibilities and drawbacks of detecting RNA biomarkers in biological fluids to improve disease-specific Accepted 2 May 2018 survival and quality of life. In particular, we examine literature on long non-coding RNAs in blood and Available online 23 May 2018 urine of prostate cancer patients. We thereby specifically focus on the need for standard operation procedures on many different levels, analytical validation, clinical validation, and assessment of clinical utility. We argue that thorough multi-step validation of putative biomarkers is necessary for successful translation into clinical prostate cancer care.
    [Show full text]
  • Normalization of Cell Associated Antiretroviral Drug Concentrations
    www.nature.com/scientificreports OPEN Normalization of cell associated antiretroviral drug concentrations with a novel RPP30 droplet digital Received: 3 April 2017 Accepted: 13 February 2018 PCR assay Published: xx xx xxxx Shetty Ravi Dyavar1, Zhen Ye2, Siddappa N. Byrareddy3, Kimberly K. Scarsi1, Lee C. Winchester1, Jonathan A. Weinhold1, Courtney V. Fletcher1 & Anthony T. Podany1 Quantifcation of antiretroviral (ARV) drug concentrations in peripheral blood mononuclear cells (PBMCs) and tissue isolated mononuclear cells (TIMCs) from lymph node (LNMC) and rectum (RMC) is an important measure of bio-distribution. Normalization of drug concentrations is critical to represent tissue drug concentrations and to analyze both intra-individual and inter-individual variability in drug distribution. However, a molecular method to normalize intracellular drug concentrations in PBMCs and TIMCs methanol extracts is currently unavailable. In this study, a novel droplet digital PCR (ddPCR) assay was designed to amplify RPP30 gene sequence conserved in human and non-human primates (NHP). Genomic DNA (gDNA) isolated from 70 percent methanol embedded PBMCs and TIMCs was used as ddPCR template to quantitate precise RPP30 copies to derive cell counts. The novel molecular method quantitated RPP30 copies in human and rhesus macaque gDNA templates with greater accuracy and precision than qPCR. RPP30 ddPCR derived cell counts are strongly correlated with automated cytometer based cell counts in PBMC (R = 0.90, p = 0.001 and n = 20); LNMC (R = 0.85 p = 0.0001 and n = 22) and RMC (R = 0.92, p = 0.0001 and n = 20) and achieved comparable normalized drug concentrations. Therefore, the RPP30 ddPCR assay is an important normalization method in drug bio-distribution and pharmacokinetic studies in humans and NHPs.
    [Show full text]
  • Master Document Template
    Copyright by Arun Kumar Ramani 2005 The Dissertation Committee for Arun Kumar Ramani Certifies that this is the approved version of the following dissertation: Towards a comprehensive human protein-protein interaction network Committee: Edward M. Marcotte, Supervisor Dean R. Appling George Georgiou David W. Hoffman Vishwanath R. Iyer Towards a comprehensive human protein-protein interaction network by Arun Kumar Ramani, B.Tech Dissertation Presented to the Faculty of the Graduate School of The University of Texas at Austin in Partial Fulfillment of the Requirements for the Degree of Doctor of Philosophy The University of Texas at Austin May, 2005 Dedication In memory of my brother Anand. Acknowledgements I would like to thank Edward Marcotte for his patience in guiding me over the past four years. I would like to acknowledge all the people in the Marcotte Lab for creating a wonderful work atmosphere and for their timely suggestions. I would like to thank my friends for their support and discussions. v Towards a comprehensive human protein-protein interaction network Publication No._____________ Arun Kumar Ramani, Ph.D. The University of Texas at Austin, 2005 Supervisor: Edward M. Marcotte Obtaining a reliable interaction data set describing the human interactome is a milestone yet to be reached. The past few years has seen tremendous progress in elucidating the yeast interactome. Experimental approaches for obtaining large-scale protein interaction data coupled with powerful computational methods for combining these data sets and for predicting functional relations between genes have been successful in tackling the yeast interactome. The concerted development of visualization techniques and the progress in the field of network biology has provided us with tools to evaluate, analyze, and interpret the interactome.
    [Show full text]
  • Modeling Microcephaly Caused by Inactivation of the Minor
    University of Connecticut OpenCommons@UConn Doctoral Dissertations University of Connecticut Graduate School 3-28-2019 Modeling Microcephaly Caused by Inactivation of the Minor Spliceosome Using a U11 Conditional Knockout Mouse Mary Baumgartner University of Connecticut - Storrs, [email protected] Follow this and additional works at: https://opencommons.uconn.edu/dissertations Recommended Citation Baumgartner, Mary, "Modeling Microcephaly Caused by Inactivation of the Minor Spliceosome Using a U11 Conditional Knockout Mouse" (2019). Doctoral Dissertations. 2076. https://opencommons.uconn.edu/dissertations/2076 ABSTRACT Modeling Microcephaly Caused by Inactivation of the Minor Spliceosome Using a U11 Conditional Knockout Mouse Mary Baumgartner, Ph.D University of Connecticut, 2019 The minor spliceosome is one of two pre-mRNA splicing machineries required for protein-coding gene expression. In multiple eukaryotic lineages, the minor spliceosome is required to remove <1% of introns (minor introns) from the pre-mRNA transcripts of minor intron-containing genes (MIGs). Despite the few minor introns in the genome, disruption of minor splicing impairs development in numerous species. In humans, minor spliceosome mutations are linked numerous developmental diseases that impact central nervous system (CNS) development. However, the role of minor splicing and MIG expression in mammalian CNS development and function remains unexplored. To address this, I characterized the expression of the minor spliceosome-specific small nuclear RNA (snRNA) components and MIGs in the developing mouse embryo and retina. This approach revealed enriched expression of the minor spliceosome-specific snRNAs in the developing mouse CNS and limb buds, and in differentiating retinal neurons. I next sought to inactivate the minor spliceosome in the developing mouse cortex (pallium) by ablating Rnu11, which encodes the minor spliceosome-specific U11 snRNA.
    [Show full text]
  • LENGTH TRANSCRIPTOME SEQUENCING a Dissertation
    HEMATOPOIETIC CELL POPULATION SEGREGATION THROUGH FULL- LENGTH TRANSCRIPTOME SEQUENCING A Dissertation submitted to the Faculty of the Graduate School of Arts and Sciences of Georgetown University in partial fulfillment of the requirements for the degree of Doctor of Philosophy in Tumor Biology By Anne Deslattes Mays, M.Sc., M.Sc. Washington, DC July 17, 2015 Copyright 2015 by Anne Deslattes Mays All Rights Reserved ii HEMATOPOIETIC CELL POPULATION SEGREGATION THROUGH FULL- LENGTH TRANSCRIPTOME SEQUENCING Anne Deslattes Mays, M.Sc., M.Sc. Thesis Advisor: Anton Wellstein, M.D. Ph.D. ABSTRACT “Progress in science results from new technologies, new discoveries and new ideas, probably in that order.” Nobel Laureate Sydney Brenner (1927 - ) Sequencing the human genome was a critical first step in setting the groundwork to understanding the molecular programming that is involved in transforming a cell from a healthy to a cancerous state. Cellular transcriptome complexity has become increasingly more apparent as technological advances have exposed us to its diversity. Full-length RNA sequencing is crucial for an unbiased analysis of transcriptome complexity. This complexity is due to posttranscriptional processing of primary transcripts that results in a variety of isoforms generated from the same genomic loci. Distinct cell lineages are defined by their transcript isoform expression profiles, and the annotation of cells can be derived from the expression of transcript isoforms that can result in functionally different proteins. Alternate splice site utilization provides cells with a powerful regulatory mechanism of gene expression that can impact the composition of the protein product, and influence the rate of translation of transcripts from multi-exon genes.
    [Show full text]
  • A Comprehensive Library of Human Transcription Factors for Cell Fate Engineering
    RESOURCE https://doi.org/10.1038/s41587-020-0742-6 A comprehensive library of human transcription factors for cell fate engineering Alex H. M. Ng1,2,3,14, Parastoo Khoshakhlagh1,2,3,14, Jesus Eduardo Rojo Arias 4,13, Giovanni Pasquini4, Kai Wang 5,6, Anka Swiersy4, Seth L. Shipman7, Evan Appleton 1,2,3, Kiavash Kiaee1,2,3, Richie E. Kohman1,2, Andyna Vernet2, Matthew Dysart1,2, Kathleen Leeper 1,2, Wren Saylor 1,2, Jeremy Y. Huang 1,2, Amanda Graveline2, Jussi Taipale8,9,10, David E. Hill 1,11, Marc Vidal1,11, Juan M. Melero-Martin5,6, Volker Busskamp 4,12 ✉ and George M. Church 1,2,3 ✉ Human pluripotent stem cells (hPSCs) offer an unprecedented opportunity to model diverse cell types and tissues. To enable systematic exploration of the programming landscape mediated by transcription factors (TFs), we present the Human TFome, a comprehensive library containing 1,564 TF genes and 1,732 TF splice isoforms. By screening the library in three hPSC lines, we discovered 290 TFs, including 241 that were previously unreported, that induce differentiation in 4 days without alteration of external soluble or biomechanical cues. We used four of the hits to program hPSCs into neurons, fibroblasts, oligodendro- cytes and vascular endothelial-like cells that have molecular and functional similarity to primary cells. Our cell-autonomous approach enabled parallel programming of hPSCs into multiple cell types simultaneously. We also demonstrated orthogonal programming by including oligodendrocyte-inducible hPSCs with unmodified hPSCs to generate cerebral organoids, which expedited in situ myelination. Large-scale combinatorial screening of the Human TFome will complement other strategies for cell engineering based on developmental biology and computational systems biology.
    [Show full text]