Integration of Large Datasets for Plant Model Organisms Yves Sucaet Iowa State University
Total Page:16
File Type:pdf, Size:1020Kb
Iowa State University Capstones, Theses and Graduate Theses and Dissertations Dissertations 2013 Integration of large datasets for plant model organisms Yves Sucaet Iowa State University Follow this and additional works at: https://lib.dr.iastate.edu/etd Part of the Bioinformatics Commons Recommended Citation Sucaet, Yves, "Integration of large datasets for plant model organisms" (2013). Graduate Theses and Dissertations. 13483. https://lib.dr.iastate.edu/etd/13483 This Dissertation is brought to you for free and open access by the Iowa State University Capstones, Theses and Dissertations at Iowa State University Digital Repository. It has been accepted for inclusion in Graduate Theses and Dissertations by an authorized administrator of Iowa State University Digital Repository. For more information, please contact [email protected]. Integration of large datasets for plant model organisms by Yves Sucaet A dissertation submitted to the graduate faculty in partial fulfillment of the requirements for the degree of DOCTOR OF PHILOSOPHY Major: Bioinformatics and Computational Biology Program of Study Committee: Eve Syrkin Wurtele, Co-Major Professor Julie A. Dickerson, Co-Major Professor Basil Nikolau Shashi K. Gadia David Fernandez-Baca Iowa State University Ames, Iowa 2013 Copyright © Yves Sucaet, 2013. All rights reserved. ii DEDICATION This work is dedicated to my daughter, Lily Claire Sucaet. iii TABLE OF CONTENTS Page DEDICATION .......................................................................................................... ii LIST OF FIGURES ................................................................................................... v LIST OF TABLES .................................................................................................... vi ACKNOWLEDGEMENTS ...................................................................................... vii ABSTRACT .............................................................................................................. ix CHAPTER 1 PLANT PATHWAY RESOURCES AND DATABASES ............ 1 Introduction and background .............................................................................. 1 The pathway database landscape ......................................................................... 4 Pathway visualization tools ................................................................................. 15 Pathway database evolution through integration................................................. 16 Applications of pathway database integration .................................................... 19 Non-plant references and opportunities for the future ........................................ 21 A survey of integrated pathway databases and tools ........................................... 23 Pathway database maintenance – an easily overlooked detail ............................ 28 Conclusion ........................................................................................................ 30 References ........................................................................................................ 31 CHAPTER 2 METNET ONLINE ........................................................................ 44 Introduction and background .............................................................................. 45 Results and implementation ................................................................................ 46 Conclusion ........................................................................................................ 54 References ........................................................................................................ 55 CHAPTER 3 METNETAPI ................................................................................. 57 Introduction and background .............................................................................. 58 Implementation considerations ........................................................................... 60 Results and deployment ...................................................................................... 64 Discussion ........................................................................................................ 73 Conclusion ........................................................................................................ 74 References ........................................................................................................ 75 iv CHAPTER 4 CANSTOREX ................................................................................ 79 Introduction and background .............................................................................. 79 Materials and methods ........................................................................................ 84 Results and implementation ................................................................................ 86 Discussion ........................................................................................................ 88 Conclusion ........................................................................................................ 90 References ........................................................................................................ 91 v LIST OF FIGURES Page Figure 1 Pathway resources with plants and humans .............................................. 5 Figure 2 A comparison of genomes sequenced for mammals and higher plants .... 6 Figure 3 The MetNet Online portal start page ......................................................... 47 Figure 4 Different browsing options ........................................................................ 49 Figure 5 A custom ethylene-related network ........................................................... 53 Figure 6 Interconnectivity between the API’s classes ............................................. 67 Figure 7 Details of a map that illustrates shared genes between pathway ............... 70 Figure 8 Cytoscape plugin developed with MetNetAPI .......................................... 71 vi LIST OF TABLES Page Table 1 Overview of plant species specific metabolic pathway databases ............ 7 Table 2 Overview of Good Databasing Practices .................................................. 19 vii ACKNOWLEDGEMENTS It is a pleasure to express my gratitude toward the many people who made this thesis possible. I'd like to thank my parents first, André Sucaet and Juliette Wynant, who taught me that I can be anything I want to be, if I'm willing to work hard enough for it. Nil volentibus arduum. I am thankful to my supervisor, Dr. Eve Syrkin Wurtele, who supported me in a number of ways. Her encouragement, sound advice, guidance and lots of good ideas enabled me to develop and grow throughout my stay at Iowa State University. Her enthusiasm, inspiration, and great efforts to explain things clearly and simply, helped to make bioinformatics fun for me. Thanks to my co-major professor Dr. Julie A. Dickerson. I would have been lost without her. I am indebted to my remaining committee members, Dr. Basil Nikolau, Dr. Sashi Gadia and Dr. David Fernandez-Bacha. Their kind assistance, wise advice, and help with various applications, has been invaluable. Special thanks go to Dr. Michael Stewart and Dr. Zhijun Wu, for influencing and helping to shape my thoughts early on in my graduate career. I am very appreciative of my many student colleagues for providing a stimulating environment in which to learn and grow. I am especially grateful to Matthew Moscou, Misha Rajaram, Fadi Towfic, Michael Zimmermann, Preeti Bais, Xiaoyoung Sun, Sean Mooney, Xinyuan Zhao and many others. I would like to express my deepest gratitude to my wife, Tamera Elaine Sucaet, for taking a leap of faith and moving with me from viii sunny Alabama to the cold winters of Iowa. All the members of the MetNet group and the Wurtele lab have been great companions along the road, especially Heather Babka and Nick Ransom. I am grateful to the secretaries and librarians involved in the interdepartmental BCB program and GDCB department at Iowa State University, for helping the departments to run smoothly and for assisting me in many different ways. I’m particularly reminded of Trish Stauble, Constance Garnett, and Lynette McBirnie- Sprecher. Lastly, I offer my regards to all of those who supported me in any respect during the completion of the project. Many people had minor, yet important, peripheral roles: Sven Aelterman, Bea Van Nieuwenhove, Taru Deva, Doug Hanson, Venkata Kishore, Rick Van Eeden, Onno Louwen, and Junli Yamada. Support for this project was provided is in part by National Science Foundation Arabidopsis 2010 grant #052026. This work additionally supported by the National Science Foundation under Award No. EEC-0813570. ix ABSTRACT This dissertation is concerned with bioinformatics data integration. The first chapter illustrates the current state of biological pathway databases in general, and in particular, plant pathway databases. Key studies are cited to illustrate the potential benefits that may come from further research into integration methods. Different models are explored to interface with the various stakeholders of biological data repositories. A public website (http://www.metnetonline.org) was built to address the role of a bioinformatics data warehouse as a server for external third parties. A dedicated API (MetNetAPI: http://www.metnetonline.org/api) accommodates bioinformaticians (and software developers in general) who wish to