How to Cite Datasets and Link to Publications

A Digital Curation Centre ‘working level’ guide How to Cite Datasets and Link to Publications Alex Ball (DCC) and Monica Duke (DCC) Please cite as: Ball, A., & Duke, M. (2015). ‘How to Cite Datasets and Link to Publications’. DCC How-to Guides. Edinburgh: Digital Curation Centre. Available online: http://www.dcc.ac.uk/resources/how-guides Digital Curation Centre, 2015. Licensed under Creative Commons Attribution 4.0 International: http://creativecommons.org/licenses/by/4.0/ How to Cite Datasets and Link to Publications Introduction This guide will help you create links between your academic publications and the underlying datasets, so that anyone viewing the publication will be able to locate the dataset and vice versa. It provides a working knowledge of the issues and challenges involved, and of how current approaches seek to address them. This guide should interest researchers and principal investigators working on data-led research, as well as the data repositories with which they work. Why cite datasets and link mechanisms allowing authors to be open about their research while still receiving due credit; metrics used them to publications? to translate such attributions into rewards for authors 1 and their institutions; and archives ensuring that the The motivation to cite datasets arises from a recog- work is permanently available for reference and reuse.5 nition that data generated in the course of research If datasets are to be regarded as first-class records of are just as valuable to the ongoing academic discourse research, as they need to be, a similar set of control as papers and monographs. Scientific journals have systems needs to be constructed around them. traditionally supported research by disseminating A major part of this work can be achieved using knowledge in such detail that first, peer scientists could a robust citation mechanism for referencing datasets judge the strength of the conclusions based on the qual- from within traditional publications. Provided the ity of the premises and research methods employed, citation contains the name of a responsible agent, it can and second, further investigations could be based upon be used to assign due credit. By providing a globally it. In many disciplines, though, the paper alone is no unique identifier, it can be used to track the impact of longer sufficient for these purposes: the underlying a particular dataset. A citation is also an ideal place to data also need to be shared.2,3,4 provide the information needed to locate and access As a medium, the journal paper owes its success the dataset. In this way, datasets can take advantage in part to the control systems put in place around it: of the infrastructure already in place to manage journal papers. 1 The term ‘dataset’ is used throughout this guide to mean a The rise of electronic journals has led to new and logically complete set of data; some systems or services prefer the terms ‘data product’ or ‘data package’. valuable services being layered over the top of papers, 2 Stodden, V. (2009). Enabling reproducible research: Open licens- among them the provision of forward links to papers ing for scientific innovation. International Journal of Communications citing the current one. Such links help the reader Law and Policy, 13, 1–25. Retrieved 2 September 2010, from http: to gauge the impact of the paper, place it within the //www.ijclp.net/files/ijclp_web-doc_1-13-2009.pdf. 3 Open to all? Case studies of openness in research. (2010, Septem- literature and in some cases gain awareness of flaws ber). Research Information Network and National Endowment for or issues discovered by others. Forward links from Science, Technology and the Arts. Retrieved 1 May 2011, from datasets to the papers that cite them provide all the http://www.rin.ac.uk/system/files/attachments/NESTA same benefits, as well as ensuring that documentation _ _ _ _ -RIN Open Science V01 0.pdf. for the dataset can be found. 4 Lynch, C. (2009). Jim Gray’s fourth paradigm and the construction of the scientific record. In T. Hey, S. Tansley & K. Tolle (Eds.), The Ultimately, bibliographic links between datasets fourth paradigm: Data-intensive scientific discovery (pp. 177–183). Redmond, WA: Microsoft Research. Retrieved 14 July 2010, from 5 Mackenzie Owen, J. (2007). The scientific article in the age of di- http://research .microsoft .com/en -us/collaboration/ gitization (ch. 2). Information Science and Knowledge Management. fourthparadigm/. Dordrecht: Springer. doi:10.1007/1-4020-5340-1. 2 and papers are a necessary step if the culture of the relies upon data, the corresponding data should be scientific and research community as a whole is to cited. shift towards data sharing, increasing the rapidity and transparency with which science advances. Unique Identification A data citation should include a persistent method for identification that is machine actionable, globally Principles of data citation unique, and widely used by a community. The FORCE11 Data Citation Synthesis Group – whose Access members include representatives of the Research Data Data citations should facilitate access to the data them- Alliance, the ICSU World Data System, and a range selves and to such associated metadata, documentation, of projects6 – has published a set of data citation code, and other materials, as are necessary for both principles.7 The principles build on earlier work in this humans and machines to make informed use of the area, most notably by CODATA,8 the (US) National referenced data. Academies of Sciences, Engineering, and Medicine,9 the DCC,10 and the Institute for Quantitative Social Persistence Science, Harvard University,11 and have been widely Unique identifiers, and metadata describing the data, endorsed.12 and its disposition, should persist – even beyond the lifespan of the data they describe. Importance Specificity and Verifiability Data should be considered legitimate, citable products Data citations should facilitate identification of, access of research. Data citations should be accorded the to, and verification of the specific data that support a same importance in the scholarly record as citations of claim. Citations or citation metadata should include other research objects, such as publications. information about provenance and fixity sufficient to facilitate verfiying that the specific timeslice, version Credit and Attribution and/or granular portion of data retrieved subsequently Data citations should facilitate giving scholarly credit is the same as was originally cited. and normative and legal attribution to all contributors to the data, recognizing that a single style or mechanism Interoperability and Flexibility of attribution may not be applicable to all data. Data citation methods should be sufficiently flexible to accommodate the variant practices among com- Evidence munities, but should not differ so much that they In scholarly literature, whenever and wherever a claim compromise interoperability of data citation practices across communities. 6 FORCE11 Data Citation Synthesis Group, URL: https://www. force11.org/datacitation/workinggroup. 7 FORCE11, Data Citation Synthesis Group. (2014). Joint de- claration of data citation principles. Retrieved from https: Data citation for authors //www.force11.org/datacitation. 8 CODATA/ITSCI Task Force on Data Citation. (2013). Out of cite, out of mind: The current state of practice, policy and technology The first half of this guide is aimed predominantly at for data citation. Data Science Journal, 12, CIDCR1–CIDCR75. researchers. It discusses the practical business of citing doi:10.2481/dsj.OSOM13-043. datasets, such as how to construct a data citation and 9 Uhlir, P. F. (Ed.). (2012). For attribution – Developing data use it in a research paper. attribution and citation practices and standards: Summary of an international workshop. Washington, D.C.: National Academies Press. Retrieved from http://www .nap .edu/openbook .php ?record_id=13564. Ways of referencing data 10 Ball, A. & Duke, M. (2012). Data citation and linking. Edinburgh, UK: Digital Curation Centre. Retrieved from http://www.dcc.ac The usual way of referencing the data directly underly- .uk/resources/briefing-papers/introduction-curation/ ing a publication is by means of a data access statement. data-citation-and-linking. For open data, this statement should say what is 11 Altman, M. & King, G. (2007). A proposed standard for the available from which repository, and provide a URL, scholarly citation of quantitative data. D-Lib Magazine, 13(3/4). identifier or accession code to help access the data. doi:10.1045/march2007-altman 12 ‘Endorse the data citation principles’, URL: https://www. For restricted data, the statement should indicate the force11.org/datacitation/endorsements. legal or ethical reason for the restriction, and provide a 3 link to a permanent record explaining the conditions of data access statements is given by some institutions.16 access. It is also possible to reference data from non- While a simple statement of this sort fulfils the basic textual outputs, such as other datasets. Indeed, doing need to reference data, it falls short in several respects: so may help satisfy the licensing conditions of the earlier datasets and encourage data sharing by supporting • if there is a typographical error in the identifier transitive credit models.17 One straightforward way or URL, there is no additional information to of doing this is to include with the dataset a table that locate the data among the repository’s holdings; lists the source datasets and indicates the subset that • authors may be tempted to give the URL of each one contributed. An early example of this was the repository, rather than one specific to the published as part of the supplementary data of a 2011 dataset; paper on microattribution for genetic variation data.18 • it does not give due credit to the creators of the Another solution is to see if the repository holding the dataset – an especially important point if these data will record the information among the metadata it are different from the authors of the publication; holds for the dataset.

How to Cite Datasets and Link to Publications

CHAPTER 12 Making Your Web Site Mashable

Life Sciences and the Web: a New Era for Collaboration

Quality Spine Care

Gigabyte: Publishing at the Speed of Research Scott C

Elie Bursztein, Baptiste Gourdin, John Mitchell Stanford University & LSV-ENS Cachan

Open Science and the Role of Publishers in Reproducible Research

The Impact of Research Data Sharing and Reuse on Data Citation in STEM Fields

ERC Document on Open Research Data and Data Management Plans

Open Science Toolbox

Assessing the Impact and Quality of Research Data Using Altmetrics and Other Indicators

Gold Open Access in the Uk: Springer Nature's Transition

Data Journals: a Survey