Proteomic Strategies for Cultural Heritage: from Bones to Paintings☆

Total Page:16

File Type:pdf, Size:1020Kb

Proteomic Strategies for Cultural Heritage: from Bones to Paintings☆ Microchemical Journal 126 (2016) 341–348 Contents lists available at ScienceDirect Microchemical Journal journal homepage: www.elsevier.com/locate/microc Proteomic strategies for cultural heritage: From bones to paintings☆ Roberto Vinciguerra a, Addolorata De Chiaro a, Piero Pucci a, Gennaro Marino b, Leila Birolo a,⁎ a Dipartimento di Scienze Chimiche, Università di Napoli Federico II, Complesso Universitario Monte S.Angelo, Via Cinthia, Napoli, Italy b Università degli Studi ‘Suor Orsola Benincasa’, 80132 Napoli, Italy article info abstract Article history: In recent years, proteomics procedures have become increasingly popular for the characterization of proteina- Received 30 July 2015 ceous materials in ancient samples of several cultural heritage objects. The knowledge of the materials used in Received in revised form 18 December 2015 a work of art is crucial, not only to give an insight in the historical context of objects and artists, but also to analyse Accepted 18 December 2015 degradation processes taking place in aged objects and to develop appropriate conservation and/or restoration Available online 29 December 2015 treatments. However, protocols routinely applied for typical modern samples still need to be fully adapted to Keywords: take into account the low amount of proteinaceous material, the heterogeneity and the unusual physical state Proteomics of the samples, as well as the high levels of damage found in ancient samples. This paper deals with some exam- Ancient proteins ples of the adaptation of classical proteomic strategies in the analysis of ancient samples to meet the different LC-MS/MS aims in the cultural heritage field. © 2015 Elsevier B.V. All rights reserved. 1. Introduction instruments are therefore perfectly adequate to afford the minimal quantities of ancient samples. However, all the steps of the proteomic Proteomics is typically considered to be associated with the study of procedure need to be thoughtfully adapted, from the optimization of living organisms; however, its inherently multidisciplinary nature has specific protocols for sample preparation to the development of data recently led to the application of proteomic methods to oddly assorted analysis tools that can cope with ancient, damaged samples. areas ranging from forensics, food analysis, clinical medicine and even Moreover, although for merely identifying purposes, the detection of for studying the origins of life on earth; it has been proven to be an as few as two peptides is sufficient to properly pinpoint the protein, the effective tool also for the scientific analysis of artworks [1]. characterization of the modifications induced by ageing and deteriora- Proteomics for cultural heritage, i.e. the identification of proteina- tion processes requires a deeper examination of most of the protein ceous material used by artists in their masterpieces and found in archae- primary structure. ological remains, is still in its infancy, with the first paper dating back to This paper deals with the reasoning behind the choice of the steps to the early 2000 [2]. In particular, protocols routinely applied for typical be carried out in the analysis of ancient samples and with the adaptation modern samples still need to be fully adapted to take into account the of classical protocols to meet the different aims. It is not intended to low amount of proteinaceous material, the heterogeneity and the cover the whole panel of possibilities and tricks but, rather to illustrate unusual physical state of the samples, as well as the high levels of dam- the logic that might address the choice of sample treatment and data age found in ancient samples. analysis to get the most of the experiment. Although analyses in proteomics are per se invasive, modern mass spectrometry instrumentations enable the characterization of proteins 2. Material and methods with extremely high sensitivity even in crude mixtures in which the dynamic range of components abundance exceeds 1000-fold, and on Ammonium hydrogen carbonate (Ambic), Ethylenediaminetetra- very limited amount of sample, typically less than 10 μg. Modern acetic acid (EDTA); Tri(hydroxymethyl)aminomethane (TRIS), Urea, GuHCl and TPCK-treated trypsin were from Sigma; Formic acid and ☆ Selected papers presented at TECHNART 2015 Conference, Catania (Italy), April 27– Acetonitrile (ACN) were purchased from Baker. Deionized water was 30, 2015. obtained from Millipore cartridge equipment. Hydrochloric acid was ⁎ Corresponding author at: Dipartimento di Scienze Chimiche, Università di Napoli, purchased from Carlo Erba. Federico II, Napoli, Italy. Models of paint layers were prepared with milk as binders and azur- E-mail addresses: [email protected] (R. Vinciguerra), [email protected] (A. De Chiaro), [email protected] (P. Pucci), [email protected] ite (Cu3(CO3)2(OH)2), minium (Pb3O4), calcite (CaCO3), and vermilion (G. Marino), [email protected] (L. Birolo). (HgS) as pigments on glass slides and skimmed milk as control without http://dx.doi.org/10.1016/j.microc.2015.12.024 0026-265X/© 2015 Elsevier B.V. All rights reserved. 342 R. Vinciguerra et al. / Microchemical Journal 126 (2016) 341–348 pigments. Paint replicas were left to dry at RT on the bench for one A licenced version of Mascot software (www.matrixscience.com) month. Bone sample was a fragment from human bone of the first version 2.4.0. was used with trypsin as enzyme; 3, as allowed num- century A.D. ber of missed cleavage; 10 ppm MS tolerance and 0.6 Da MS/MS tol- erance; peptide charge from +2 to +3. No fixed chemical fi 2.1 Sample treatment modi cation was inserted, but possible oxidation of methionines, formation of pyroglutamic acid from glutamine residues at the N- Urea pre-treatment: 10 μL of a solution of 6 M Urea was added to terminal position of peptides, and deamidation at asparagines and fi micro-samples (ca 300–800 μg) and incubated for 10 min at RT, follow- glutamines were considered as variable modi cations [4].When fi fi ed by sonication for 20 min. Urea was then 6-fold diluted with water. collagen proteins were identi ed, a further identi cation run was EDTA pre-treatment: about 100 μL of a solution of 0.5 M EDTA was carried out, with the insertion of hydroxylation on lysine and proline fi fi fi added to the bone fragment for 10 days at RT, refreshing the solution as variable modi cations, since more con dent identi cations are every 2 days. After centrifugation for 2 min at 10,000 rpm in a benchtop commonly obtained for these proteins by taking into consideration fi microfuge, the Urea protocol described above was applied. their extensive post-translational modi cations [4].Onlyproteins HCl pre-treatment: 50 μL of 0.6 M HCl was added to the bone frag- presenting two or more peptides were considered as positively iden- fi ment and incubated at 4 °C for 4 h. After centrifugation for 2 min at ti ed. Individual ion score threshold provided by Mascot software to 10,000 rpm in a benchtop microfuge the supernatant was removed evaluate the quality of matches in MS/MS data was generally 31 for and washed with 20 μL of 10 mM Ambic. Washes were repeated for paintings and 43 for human samples. Spectra with Mascot score b four times. 100 μL of 50 mM Ambic was added and sample was left at 10 were rejected. 65 °C for 3 h. After centrifugation at 10,000 rpm for 15 min the superna- tant was removed. 3. Results and discussion GuHCl pre-treatment: 200 μL of 0.6 M HCl was added and incubated at 4 °C for 18 h. After centrifugation for 1 min at 14,000 rpm in a bench- 3.1 The minimal protocol top microfuge, and the acid-insoluble pellet washed three times with “ ” 200 μL of distilled water. The pellet was incubated at 4 °C for 72 h in a The commonly used bottom-up approach to identify proteins is buffer containing 100 mM Tris and 6 M GuHCl at pH 7.4. The sample based on the enzymatic digestion of the proteins and can be directly was then centrifuged for 1 min at 14,000 rpm in a benchtop microfuge. performed in heterogeneous phase on a sample fragment followed by The supernatant was buffer-exchanged into 10 mM Ambic using 3 K mass spectrometric analyses of the released peptides (Fig. 1). Even molecular weight cut-off Amicon Ultra, centrifugal filter unit. when the protein is embedded in a complex mixture such as that of a painting layer, few peptides released by the protease without any pre- treatment of the sample are enough to identify the protein [3], without 2.2 Protein digestion and LC-MS/MS analysis significantly affecting the sample itself. This avoid the extraction of the whole protein from the sample using harsh methods, while digestion After any pre-treatment of the sample, enzymatic digestion was can rather be carried out by depositing on the surface of a small sample carried out as in the minimally invasive proteomic analytical procedure an aqueous neutral solution containing the enzyme that directly trims fl fi described by Leo et al. [3].Brie y, trypsin was added to a nal concentra- protruding peptides. The solution will then be gently removed and, μ – μ tion of 10 ng/ L to micro-samples (ca 300 800 g) as directly suspended once released, peptides can be analysed by mass spectrometric methods μ in 50 L of Ambic 10 mM. After incubation at 37 °C for 16 h, the superna- such as MALDI-TOF and LC-MS/MS, and database searches by bioinfor- tants were recovered by centrifugation at 10,000 rpm, and the peptide matics tools such as Mascot (www.matrixscience.com) allow protein fi μ mixture was ltered on 0.22 m PVDF membrane (Millipore), concen- identification.
Recommended publications
  • Ancient Metaproteomics: a Novel Approach for Understanding Disease And
    Ancient metaproteomics: a novel approach for understanding disease and diet in the archaeological record Jessica Hendy PhD University of York Archaeology August, 2015 ii Abstract Proteomics is increasingly being applied to archaeological samples following technological developments in mass spectrometry. This thesis explores how these developments may contribute to the characterisation of disease and diet in the archaeological record. This thesis has a three-fold aim; a) to evaluate the potential of shotgun proteomics as a method for characterising ancient disease, b) to develop the metaproteomic analysis of dental calculus as a tool for understanding both ancient oral health and patterns of individual food consumption and c) to apply these methodological developments to understanding individual lifeways of people enslaved during the 19th century transatlantic slave trade. This thesis demonstrates that ancient metaproteomics can be a powerful tool for identifying microorganisms in the archaeological record, characterising the functional profile of ancient proteomes and accessing individual patterns of food consumption with high taxonomic specificity. In particular, analysis of dental calculus may be an extremely valuable tool for understanding the aetiology of past oral diseases. Results of this study highlight the value of revisiting previous studies with more recent methodological approaches and demonstrate that biomolecular preservation can have a significant impact on the effectiveness of ancient proteins as an archaeological tool for this characterisation. Using the approaches developed in this study we have the opportunity to increase the visibility of past diseases and their aetiology, as well as develop a richer understanding of individual lifeways through the production of molecular life histories. iii iv List of Contents Abstract ...............................................................................................................................
    [Show full text]
  • Decelerated Genome Evolution in Modern Vertebrates Revealed by Analysis of Multiple Lancelet Genomes
    ARTICLE Received 20 May 2014 | Accepted 18 Nov 2014 | Published 19 Dec 2014 DOI: 10.1038/ncomms6896 OPEN Decelerated genome evolution in modern vertebrates revealed by analysis of multiple lancelet genomes Shengfeng Huang1, Zelin Chen1, Xinyu Yan1, Ting Yu1, Guangrui Huang1, Qingyu Yan1, Pierre Antoine Pontarotti2, Hongchen Zhao1, Jie Li1, Ping Yang1, Ruihua Wang1, Rui Li1, Xin Tao1, Ting Deng1, Yiquan Wang3,4, Guang Li3,4, Qiujin Zhang5, Sisi Zhou1, Leiming You1, Shaochun Yuan1, Yonggui Fu1, Fenfang Wu1, Meiling Dong1, Shangwu Chen1 & Anlong Xu1,6 Vertebrates diverged from other chordates B500 Myr ago and experienced successful innovations and adaptations, but the genomic basis underlying vertebrate origins are not fully understood. Here we suggest, through comparison with multiple lancelet (amphioxus) genomes, that ancient vertebrates experienced high rates of protein evolution, genome rearrangement and domain shuffling and that these rates greatly slowed down after the divergence of jawed and jawless vertebrates. Compared with lancelets, modern vertebrates retain, at least relatively, less protein diversity, fewer nucleotide polymorphisms, domain combinations and conserved non-coding elements (CNE). Modern vertebrates also lost substantial transposable element (TE) diversity, whereas lancelets preserve high TE diversity that includes even the long-sought RAG transposon. Lancelets also exhibit rapid gene turnover, pervasive transcription, fastest exon shuffling in metazoans and substantial TE methylation not observed in other invertebrates. These new lancelet genome sequences provide new insights into the chordate ancestral state and the vertebrate evolution. 1 State Key Laboratory of Biocontrol, Guangdong Key Laboratory of Pharmaceutical Functional Genes, School of Life Sciences, Sun Yat-sen University, Guangzhou 510275, China. 2 Evolution Biologique et Mode´lisation UMR 7353 Aix Marseille Universite´/CNRS, 3 Place Victor Hugo, 13331 Marseille, France.
    [Show full text]
  • Analysis of 5000 Year-Old Human Teeth Using Optimized Large-Scale And
    Analysis of 5000 year-old human teeth using optimized large-scale and targeted proteomics approaches for detection of sex-specific peptides Carine Froment, Mathilde Hourset, Nancy Sáenz-Oyhéréguy, Emmanuelle Mouton-Barbosa, Claire Willmann, Clément Zanolli, Rémi Esclassan, Richard Donat, Catherine Thèves, Odile Burlet-Schiltz, et al. To cite this version: Carine Froment, Mathilde Hourset, Nancy Sáenz-Oyhéréguy, Emmanuelle Mouton-Barbosa, Claire Willmann, et al.. Analysis of 5000 year-old human teeth using optimized large-scale and targeted proteomics approaches for detection of sex-specific peptides. Journal of Proteomics, Elsevier, 2020, 211, pp.103548. 10.1016/j.jprot.2019.103548. hal-02322441 HAL Id: hal-02322441 https://hal.archives-ouvertes.fr/hal-02322441 Submitted on 18 Nov 2020 HAL is a multi-disciplinary open access L’archive ouverte pluridisciplinaire HAL, est archive for the deposit and dissemination of sci- destinée au dépôt et à la diffusion de documents entific research documents, whether they are pub- scientifiques de niveau recherche, publiés ou non, lished or not. The documents may come from émanant des établissements d’enseignement et de teaching and research institutions in France or recherche français ou étrangers, des laboratoires abroad, or from public or private research centers. publics ou privés. Journal Pre-proof Analysis of 5000year-old human teeth using optimized large- scale and targeted proteomics approaches for detection of sex- specific peptides Carine Froment, Mathilde Hourset, Nancy Sáenz-Oyhéréguy, Emmanuelle Mouton-Barbosa, Claire Willmann, Clément Zanolli, Rémi Esclassan, Richard Donat, Catherine Thèves, Odile Burlet- Schiltz, Catherine Mollereau PII: S1874-3919(19)30320-3 DOI: https://doi.org/10.1016/j.jprot.2019.103548 Reference: JPROT 103548 To appear in: Journal of Proteomics Received date: 24 January 2019 Revised date: 30 August 2019 Accepted date: 7 October 2019 Please cite this article as: C.
    [Show full text]
  • SYSTEMS BIOLOGY Ancient Protein Complexes Revealed
    RESEARCH HIGHLIGHTS SYSTEMS BIOLOGY Ancient protein complexes revealed A systematic analysis of protein-complex history, they applied the approach to eight some interesting conclusions. For example, composition across a billion years of evo- species—mouse, frog, fly, worm, sea urchin, despite the fact that the majority of human lution reveals a spectrum of conservation. sea anemone, amoeba and yeast—spanning proteins arose in animals, they found that by Comparative analyses of protein-protein a billion years of evolution. and large, most stable human protein com- interaction networks from Escherichia coli to With this massive set of experiments, the plexes were likely inherited from a unicellular humans have shown that such networks are researchers identified and quantified 13,386 ancestor. Membership in such ancient com- conserved and relatively slow-evolving. Yet protein orthologues. They developed a plexes was either conserved or mixed with open questions remain about whether physi- machine learning–classification method to animal-specific components, and these com- cal interactions are preserved between evolu- pull out interactions that were conserved plexes tended to be abundant, ubiquitously tionarily distinct species, what these protein between human proteins and their worm, expressed and functionally associated with complexes are and what they do. fly, mouse and sea urchin orthologues (the core cellular processes. The relatively rare These are some of the questions that data from the other four species were used complexes consisting solely of proteins that Andrew Emili of the University of Toronto, to benchmark the approach). Their final net- arose in animals, in contrast, tended to have Edward Marcotte of the University of Texas at work contained 16,655 human co-complex functions strongly linked to multicellularity.
    [Show full text]
  • Ancient Protein Analysis in Archaeology
    This is a repository copy of Ancient protein analysis in archaeology. White Rose Research Online URL for this paper: https://eprints.whiterose.ac.uk/171551/ Version: Published Version Article: Hendy, Jessica orcid.org/0000-0002-3718-1058 (2021) Ancient protein analysis in archaeology. Science Advances. eabb9314. ISSN 2375-2548 https://doi.org/10.1126/sciadv.abb9314 Reuse This article is distributed under the terms of the Creative Commons Attribution-NonCommercial (CC BY-NC) licence. This licence allows you to remix, tweak, and build upon this work non-commercially, and any new works must also acknowledge the authors and be non-commercial. You don’t have to license any derivative works on the same terms. More information and the full terms of the licence here: https://creativecommons.org/licenses/ Takedown If you consider content in White Rose Research Online to be in breach of UK law, please notify us by emailing [email protected] including the URL of the record and the reason for the withdrawal request. [email protected] https://eprints.whiterose.ac.uk/ SCIENCE ADVANCES | REVIEW ANTHROPOLOGY Copyright © 2021 The Authors, some Ancient protein analysis in archaeology rights reserved; exclusive licensee Jessica Hendy1,2 American Association for the Advancement of Science. No claim to The analysis of ancient proteins from paleontological, archeological, and historic materials is revealing insights original U.S. Government into past subsistence practices, patterns of health and disease, evolution and phylogeny, and past environments. Works. Distributed This review tracks the development of this field, discusses some of the major methodological strategies used, and under a Creative synthesizes recent developments in archeological applications of ancient protein analysis.
    [Show full text]
  • Protein Sequence Comparison of Human and Non-Human Primate
    Protein sequence comparison of human and non-human primate tooth proteomes Carine Froment, Clément Zanolli, Mathilde Hourset, Emmanuelle Mouton-Barbosa, Andreia Moreira, Odile Burlet-Schiltz, Catherine Mollereau To cite this version: Carine Froment, Clément Zanolli, Mathilde Hourset, Emmanuelle Mouton-Barbosa, Andreia Moreira, et al.. Protein sequence comparison of human and non-human primate tooth proteomes. Journal of Proteomics, Elsevier, 2021, 231, pp.104045. 10.1016/j.jprot.2020.104045. hal-03039831 HAL Id: hal-03039831 https://hal.archives-ouvertes.fr/hal-03039831 Submitted on 4 Dec 2020 HAL is a multi-disciplinary open access L’archive ouverte pluridisciplinaire HAL, est archive for the deposit and dissemination of sci- destinée au dépôt et à la diffusion de documents entific research documents, whether they are pub- scientifiques de niveau recherche, publiés ou non, lished or not. The documents may come from émanant des établissements d’enseignement et de teaching and research institutions in France or recherche français ou étrangers, des laboratoires abroad, or from public or private research centers. publics ou privés. Protein sequence comparison of human and non-human primate tooth proteomes Carine Froment1, Clément Zanolli2, Mathilde Hourset3,4, Emmanuelle Mouton-Barbosa1, Andreia Moreira3, Odile Burlet-Schiltz1 and Catherine Mollereau3 1 Institut de Pharmacologie et Biologie Structurale (IPBS), Université de Toulouse, CNRS, UPS, Toulouse, France. 2 Laboratoire PACEA, UMR 5199 CNRS, Université de Bordeaux, Pessac, France.
    [Show full text]
  • Functional Innovation from Changes in Protein Domains and Their Combinations
    Functional innovation from changes in protein domains and their combinations Jonathan G. Lees1*+, Natalie L. Dawson1, Ian Sillitoe1 and Christine A. Orengo1 1 Institute of Structural and Molecular Biology, Division of Biosciences, University College London, Gower Street, London, WC1E 6BT, UK *To whom correspondence should be addressed. Tel: 02076793890; Fax: 02076797193; Email: [email protected] Abstract Domains are the functional building blocks of proteins. In this work we discuss how domains can contribute to the evolution of new functions. Domains themselves can evolve through various mechanisms, altering their intrinsic function. Domains can also facilitate functional innovations by combining with other domains to make novel proteins. We discuss the mechanisms by which domain and domain combinations support functional innovations. We highlight interesting examples where changes in domain combination promote changes at the domain level. Introduction Globular protein domains are structurally compact, independently folding units and can be grouped into sometimes very large homologous superfamilies (Sillitoe et al. 2015). Domain superfamilies can be defined using purely sequence (as in Pfam (Finn et al. 2015)), or sequence combined with structural data (such as used in the CATH and SCOP resources (Sillitoe et al. 2015; Murzin et al. 1995)). As the amount of protein structure data has grown, it has become clear that some homologues can diverge significantly in their structures and functions. In some superfamilies the domain fold can change to such an extent, that it appears similar to domain folds in other superfamilies. Some recent domain structure classifications (e.g. SCOP2 (Andreeva et al. 2014) highlight these structural overlaps, which suggest continuity in some regions of structure space (Edwards & Deane 2015).
    [Show full text]
  • Did Evolution Leap to Create the Protein Universe? Burkhard Rost
    409 Did evolution leap to create the protein universe? Burkhard Rost The genomes of over 60 organisms from all three kingdoms of Although most of these analyses utilize a wealth of biological life are now entirely sequenced. In many respects, the inventory data, they are not explicitly based on the fact that we have of proteins used in different kingdoms appears surprisingly entire genome sequences from representatives of all three similar. However, eukaryotes differ from other kingdoms in that kingdoms of life: eukarya, bacteria (prokaryotes) and they use many long proteins, and have more proteins with archaea. What do we learn from generating lists of parts of coiled-coil helices and with regions abundant in regular the whole [10,11•]? secondary structure. Particular structural domains are used in many pathways. Nevertheless, one domain tends to occur only Here, I focus on findings from methods that endeavor to once in one particular pathway. Many proteins do not have capture overall features for entire organisms. I challenge close homologues in different species (orphans) and there the assumption that bioinformatics is slowly but steadily could even be folds that are specific to one species. This view approaching the point at which we can smoothly move implies that protein fold space is discrete. An alternative model through neighborhoods of protein relationships in order to suggests that structure space is continuous and that modern generate an atlas of the fates and functions of proteins in proteins evolved by aggregating fragments of ancient proteins. the context of the cell. Either way, after having harvested proteomes by applying standard tools, the challenge now seems to be to develop Overview of proteomes: catalogs of structure better methods for comparative proteomics.
    [Show full text]
  • Elucidation of Cross-Species Proteomic Effects in Human and Hominin Bone Proteome Identification Through a Bioinformatics Experiment F
    Welker BMC Evolutionary Biology (2018) 18:23 https://doi.org/10.1186/s12862-018-1141-1 RESEARCH ARTICLE Open Access Elucidation of cross-species proteomic effects in human and hominin bone proteome identification through a bioinformatics experiment F. Welker1,2 Abstract Background: The study of ancient protein sequences is increasingly focused on the analysis of older samples, including those of ancient hominins. The analysis of such ancient proteomes thereby potentially suffers from “cross-species proteomic effects”: the loss of peptide and protein identifications at increased evolutionary distances due to a larger number of protein sequence differences between the database sequence and the analyzed organism. Error-tolerant proteomic search algorithms should theoretically overcome this problem at both the peptide and protein level; however, this has not been demonstrated. If error-tolerant searches do not overcome the cross-species proteomic issue then there might be inherent biases in the identified proteomes. Here, a bioinformatics experiment is performed to test this using a set of modern human bone proteomes and three independent searches against sequence databases at increasing evolutionary distances: the human (0 Ma), chimpanzee (6-8 Ma) and orangutan (16-17 Ma) reference proteomes, respectively. Results: Incorrectly suggested amino acid substitutions are absent when employing adequate filtering criteria for mutable PeptideSpectrumMatches(PSMs),butroughlyhalfofthemutablePSMswerenotrecovered.Asaresult,peptideand protein identification rates are higher in error-tolerant mode compared to non-error-tolerant searches but did not recover protein identifications completely. Data indicates that peptide length and the number of mutations between the target and database sequences are the main factors influencing mutable PSM identification. Conclusions: The error-tolerant results suggest that the cross-species proteomics problem is not overcome at increasing evolutionary distances, even at the protein level.
    [Show full text]
  • Ancient Evolutionary Signals of Protein-Coding
    Casimiro-Soriguer et al. BMC Genomics (2020) 21:210 https://doi.org/10.1186/s12864-020-6632-y RESEARCH ARTICLE Open Access Ancient evolutionary signals of protein- coding sequences allow the discovery of new genes in the Drosophila melanogaster genome Carlos S. Casimiro-Soriguer, Alejandro Rubio, Juan Jimenez and Antonio J. Pérez-Pulido* Abstract Background: The current growth in DNA sequencing techniques makes of genome annotation a crucial task in the genomic era. Traditional gene finders focus on protein-coding sequences, but they are far from being exhaustive. The number of this kind of genes continuously increases due to new experimental data and development of improved bioinformatics algorithms. Results: In this context, AnABlast represents a novel in silico strategy, based on the accumulation of short evolutionary signals identified by protein sequence alignments of low score. This strategy potentially highlights protein-coding regions in genomic sequences regardless of traditional homology or translation signatures. Here, we analyze the evolutionary information that the accumulation of these short signals encloses. Using the Drosophila melanogaster genome, we stablish optimal parameters for the accurate gene prediction with AnABlast and show that this new strategy significantly contributes to add genes, exons and pseudogenes regions, yet to be discovered in both already annotated and new genomes. Conclusions: AnABlast can be freely used to analyze genomic regions of whole genomes where it contributes to complete the previous annotation. Keywords: Ancient sequences, Gene finding, Protein-coding genes, Drosophila melanogaster, Protomotifs Background significant number of protein-coding sequences and other Research groups from all over the world are sequencing functional genomic elements are missing when using cur- whole genomes as a common task, taking advantage of rently available genomic annotation approaches.
    [Show full text]
  • Markgerstein Loaded 2024 Tweets
    MarkGerstein loaded 2024 tweets Hide Replies Hide Retweets Change User Follow @felixturner 60.1K followers .@kgebo & Devaney provide a nice overview of the @AllofUsResearch program as first building relationships with part… https://t.co/KQp0wX6sgq Jun 24, 2019 Posting my talk tomorrow at the @MRC_LMB in @Cambridge_Uni https://t.co/yJWjJgykIV Addressing core Qs in neuro-geno… https://t.co/9zzaE42mCX Jun 19, 2019 Posting my talk tomorrow on cancer genomics @DBMR_UniBe (in @UniBern), hosted by @MarkARubin1… https://t.co/v15JTy4X4A Jun 13, 2019 Posting my talk later today at @IEEEembs BHI meeting in Chicago https://t.co/aZmnjmzIH8 Quantification of sensitive… https://t.co/L0fl2q1Suw May 21, 2019 Images from the end of the day @Yale - Sunset, dusk & beyond for #Yale2019. Congrats to the new graduates. (Picture… https://t.co/74fCAiK9Bi May 21, 2019 Posting my talk tomorrow in Boston in the SynGen Series https://t.co/9dRc6NUblN Finding drug targets via deep-lear… https://t.co/c9SYYM4djh May 15, 2019 At #BoG19, @BPasaniuc gives a nice overview of SNP heritability in previous literature. How it is enriched in vario… https://t.co/N7apg9tWGK May 11, 2019 For #BoG19's #GenomesSparkJoy picture compilation, I've been busy snapping shots of the great sculptures @CSHL. My… https://t.co/vTXHWdEDQd May 11, 2019 At #BoG19, @genemodeller uses PRSes for coronary risk to further stratify high- & low-risk individuals (based on tr… https://t.co/Ook1ReKBUc May 09, 2019 At #BoG19, @genemodeller starts off by describing a huge matrix of 14M SNPs X 10K phenotype & trait measurements. W… https://t.co/OofRxo2I55 May 09, 2019 At #BoG19, Ongen describes a method to find CRD (cis-regulatory domains) that may be enriched for non-coding cancer… https://t.co/TfJ0rlG42y May 09, 2019 At #BoG19, @Kasper_Lage shows a nice visualization of both the Manhattan plot from GWAS & protein-protein interacti… https://t.co/3PF1AdRLA4 May 09, 2019 At #BoG19, @Ruthie_Johnson shows that heritability is highly correlated with the number of causal SNPs.
    [Show full text]
  • Probabilistic Protein Engineering
    Probabilistic Protein Engineering Thesis by Kevin Kaichuang Yang In Partial Fulfillment of the Requirements for the Degree of Doctor of Philosophy in Chemical Engineering CALIFORNIA INSTITUTE OF TECHNOLOGY Pasadena, California 2019 Defended 14 Dec 2018 ii © 2019 Kevin Kaichuang Yang ORCID: 0000-0001-9045-6826 Some rights reserved. This thesis is distributed under a Creative Commons Attribution-NonCommercial-ShareAlike License iii ACKNOWLEDGEMENTS First, I’d like to thank my advisor Frances Arnold for her support and guidance throughout my PhD. I applied to Caltech because I wanted to work for Frances, and she has not disappointed. Even as I took a detour to teach high school, she welcomed me into her lab first for a summer, and then for my PhD. Frances has a keen eye for important but tractable scientific problems and gave me the freedom to pursue them past the boundaries of her discipline. Frances has pushed me to be a better scientist, and her lessons are the foundation for the rest of my career. One thing Frances has done very well is to build a world-class research group. I would like to specifically thank several group members. During my first stint in the lab, Sabine Brinkmann-Chen, John McIntosh, and Chris Farwell taught me to perform directed evolution. Sabine also keeps us all safe, happy, and productive. When I rejoined the lab, Lukas Herwig reacquainted me with molecular biology. It was a joy to work with Claire Bedbrook and Austin Rice on my first three papers. Thank you for collecting my data for me. Zachary Wu has been a great co-belligerent for the cause of machine learning in the Arnold lab.
    [Show full text]