Research Data Management in Canada: a Backgrounder
Total Page:16
File Type:pdf, Size:1020Kb
Research Data Management in Canada: A Backgrounder July 18, 2019 Based on a report produced for Innovation, Science and Economic Development (ISED) by the Canadian Association for Research Libraries (CARL), Consortia Advancing Standards in Research Administration Information (CASRAI), Leadership Council for Digital Research Infrastructure (LCDRI), Research Data Canada (RDC). Contributors: David Baker, Donna Bourne-Tyson, Laura Gerlitz, Susan Haigh, Shahira Khair, Mark Leggott, Jeff Moon, Chantel Ridsdale, Robbin Tourangeau, Martha Whitehead DOI: 10.5281/zenodo.3341596 1 Table of Contents Table of Contents Introduction Definitions Background to the Document The Research Lifecycle and RDM Functions Plan Create Process Analyze Disseminate Preserve Reuse Store Discover Document and Curate Secure The Continuum of Research Data Storage Cost Effectiveness Risk Mitigation Regional Talent Development The Impact of Good Research Data Management Growth in Data Production Trends Driving RDM Initiatives Journal publishers Funding Agencies Researchers The Canadian Landscape Local Provincial/Regional National Traditional Knowledge and Ways of Knowing The International Landscape 2 International Other Contributors Challenges and Opportunities Current Opportunities for RDM in Canada What are the Current Challenges for RDM in Canada? A National Vision Vision Principles Goals Selective Bibliography Canadian International United Kingdom United States 3 Introduction Digital research infrastructure is transforming the practice of research by enabling the rapid creation of massive quantities of data at an explosive rate. We are only beginning to understand its total impact on the research process, but we recognize the changes it brings fundamentally reshape the speed at which our researchers work, the questions they can ask, and the results they can achieve. To fully comprehend the social and economic potential that new digital research infrastructure and data offer to Canadians, and to ensure that we are among the first to mine the benefits of this important resource, a strong and vibrant digital research infrastructure (DRI) ecosystem must be in place. This DRI must allow Canadian researchers to store, access, reuse, and build upon digital research data that is essential to their ongoing capacity to remain current and collaborative in their fields, enabling them to generate and contribute critical research that underpins Canada’s economic and social well-being. Accessible digital research data also holds great potential benefits for the Canadian nonprofit and private sectors in helping them advance critical social and commercialization goals, though for it to be of any use it must be managed appropriately. The systematic management of research data is an essential component of the DRI ecosystem, and Canada has begun to lay the supportive foundations to enable researchers to thrive in this ecosystem through the development and delivery of high quality, equitable, and effective data management (DM) services and platforms. However, for Canada to become a global leader in data management and ensure we are continuing to produce world-class research in a changing research landscape, we must continue to build on those DM foundations. This report provides an overview of the DM environment in Canada, and identifies key challenges to achieving this goal, and will be updated annually to communicate changes in the landscape. 4 Definitions Advanced Research Computing provides researchers with digital technology and expertise to help them solve research issues that are either too large or too complex for them to undertake on their own. It includes services, advice, hardware and software, all supported by highly qualified personnel (HQP), to enable research activities with significant data or computation requirements, including data acquisition, simulation, experimentation, analysis, and exploration. It is known by a number of different names, depending on the audience. Examples include high performance computing (HPC), cyberinfrastructure, and supercomputing.1 Data2 are facts or observations captured with a minimum of contextual interpretation. Data may be in any format or medium (analog or digital). Data elements can include text (and other symbols), numbers, images (and other graphical representations), video or audio. Data includes metadata. There are two types of data pertaining to the university research environment that are the focus of this data management position paper: research data and research management information. Data Management are the activities of data policies, data planning, data element standardization, information management control, data synchronization, data sharing, and database development, including practices and projects that acquire, control, protect, deliver and enhance the value of data and information. Digital Research Infrastructure3 are those layers that sit between base technology (a computer science concern) and discipline-specific science. The focus is on value-added systems and services that can be widely shared across scientific domains, both supporting and enabling large increases in multi- and interdisciplinary science while reducing duplication of effort and resources (e.g., including hardware, software, personnel, services and organizations). In Canada, the preferred term has become Digital Infrastructure to refer to what is also known as Cyber-Infrastructure or e-Research Infrastructure. Research Data (RD) are used as primary sources to support technical or scientific enquiry, research, scholarship, or artistic activity, and as evidence in the research process and/or are 1 The Leadership Council for Digital Research Infrastructure, “Advanced Research Computing (ARC) Position Paper: For Innovation, Science, and Economic Development Canada,” The Leadership Council for Digital Research Infrastructure, August 31, 2017, 5-6. 2 The following definitions were developed through consultation with the CASRAI IRiDiuM Glossary and the Data Management (DM) Working Group (WG). They were then shared for input from LCDRI members. There is a wide and varied understanding of the terms, which can create confusion when working in a diverse and multi-disciplinary community such as Canada’s academic community. The group did its best to be inclusive and understandable in the definitions, in order to help the community to be successful in using the same shared definitions going forward. 3 This definition is taken from CASRAI's term for Digital Infrastructure; “Digital Infrastructure - CASRAI Dictionary,” CASRAI, August 13, 2015, https://dictionary.casrai.org/Digital_infrastructure. 5 commonly accepted in the research community as necessary to validate research findings and results. All other digital and non-digital content have the potential of becoming research data. Research data include metadata and may be experimental data, observational data, operational data, third-party data, public-sector data, monitoring data, processed data, or repurposed data. Research Management Information (RMI) is information used primarily to facilitate research management by research-funding and research-performing organizations. Examples include information about the people, organizations, funding, equipment, projects, outputs, outcomes and impacts of the research lifecycle. RMI include metadata recorded with the information, such as the version number of a classification scheme used to classify the expertise of a Principal Investigator or a specific project. Common synonyms for RMI include: admin data, research administration data, research information, and research documentation. Research Data Management/Data Management (RDM/DM) refers to the storage, access and preservation of data produced from a given investigation. Data management practices cover the entire lifecycle of the data, from planning the investigation to conducting it, and from backing up data as it is created and used to long term preservation of data deliverables after the research investigation has concluded. Specific activities and issues that fall within the category of data management include: File naming (the proper way to name computer files); data quality control and quality assurance; data access; data documentation (including levels of uncertainty); metadata creation and controlled vocabularies; data storage; data archiving and preservation; data sharing and reuse; data integrity; data security; data privacy; data rights; notebook protocols (lab or field). Metadata is literally "data about data". It is data that defines and describes the characteristics of other data, used to improve both business and technical understanding of data and data-related processes. Business metadata includes the names and business definitions of subject areas, entities and attributes, attribute data types and other attribute properties, range descriptions, valid domain values and their definitions. Technical metadata includes physical database table and column names, column properties, and the properties of other database objects, including how data is stored. Process metadata is data that defines and describes the characteristics of other system elements (processes, business rules, programs, jobs, tools, etc.). Data stewardship metadata is data about data stewards, stewardship processes and responsibility assignments. Background to the Document The original text for this document was based on the submission from the Leadership Council for Digital Research Infrastructure (LCDRI)