Database, 2017, 1–11 doi: 10.1093/database/baw147 Original article Original article The BioC-BioGRID corpus: full text articles anno- tated for curation of protein–protein and genetic interactions Rezarta Islamaj Dogan 1,†, Sun Kim1,†, Andrew Chatr-aryamontri2,†, Christie S. Chang3, Rose Oughtred3, Jennifer Rust3, W. John Wilbur1, Donald C. Comeau1, Kara Dolinski3 and Mike Tyers2,4 1National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, MD20894, USA, 2Institute for Research in Immunology and Cancer, Universite´de Montre´al, Canada Montre´al, QC H3C 3J7, 3Lewis-Sigler Institute for Integrative Genomics, Princeton University, Princeton, NJ 08544, USA and 4Mount Sinai Hospital, The Lunenfeld-Tanenbaum Research Institute, Canada *Corresponding author: Tel: 301 435 8769; E-mail:
[email protected] †These authors contributed equally to this work. Citation details: Islamaj Dogan,R., Kim,S., Chatr-Aryamontri,A. et al. The BioC-BioGRID corpus: full text articles annotated for curation of protein–protein and genetic interactions. Database (2016) Vol. 2016: article ID baw147; doi:10.1093/data- base/baw147 Received 30 July 2016; Revised 14 October 2016; Accepted 18 October 2016 Abstract A great deal of information on the molecular genetics and biochemistry of model organ- isms has been reported in the scientific literature. However, this data is typically described in free text form and is not readily amenable to computational analyses. To this end, the BioGRID database systematically curates the biomedical literature for gen- etic and protein interaction data. This data is provided in a standardized computationally tractable format and includes structured annotation of experimental evidence.