THE CENDARI WHITE BOOK of ARCHIVES Jakub Beneš, Natasa Bulatovic, Jennifer Edmond, Milica Knežević„ Jörg Lehmann, Francesca Morselli, Andrei Zamoiski
Total Page:16
File Type:pdf, Size:1020Kb
THE CENDARI WHITE BOOK OF ARCHIVES Jakub Beneš, Natasa Bulatovic, Jennifer Edmond, Milica Knežević„ Jörg Lehmann, Francesca Morselli, Andrei Zamoiski To cite this version: Jakub Beneš, Natasa Bulatovic, Jennifer Edmond, Milica Knežević„ Jörg Lehmann, et al.. THE CEN- DARI WHITE BOOK OF ARCHIVES : Data Exchange Recommendations for Cultural Heritage In- stitutions and Infrastructure Projects. [Research Report] Trinity College Dublin. 2016. hal-01419148 HAL Id: hal-01419148 https://hal.archives-ouvertes.fr/hal-01419148 Submitted on 18 Dec 2016 HAL is a multi-disciplinary open access L’archive ouverte pluridisciplinaire HAL, est archive for the deposit and dissemination of sci- destinée au dépôt et à la diffusion de documents entific research documents, whether they are pub- scientifiques de niveau recherche, publiés ou non, lished or not. The documents may come from émanant des établissements d’enseignement et de teaching and research institutions in France or recherche français ou étrangers, des laboratoires abroad, or from public or private research centers. publics ou privés. THE CENDARI WHITE BOOK OF ARCHIVES Data Exchange Recommendations for Cultural Heritage Institutions and Infrastructure Projects THE CENDARI WHITE BOOK OF ARCHIVES AUTHOR(S) Jakub Beneš, Nataša Bulatović, Jennifer Edmond, Milica Knežević, Jörg Lehmann, Francesca Morselli, Andrei Zamoiski COLLABORATOR(S) THEMES Archives, Data Exchange, Metadata, Repository, Data Ingestion, Cultural Heritage Institutions PERIOD 19th century - 21st century CENDARI is funded by the European Commission’s 7th Framework Programme for Research 3 TABLE OF CONTENTS 7 EXECUTIVE SUMMARY 22 BEST PRACTICES Introduction 9 INTRODUCTION Pan-European Aggregators Europeana and The European Library 10 TWO HISTORICAL REALMS, TWO WAYS OF AGGREGATING CONTENT National Aggregators Archives Hub UK 12 LEGAL FRAMEWORK National Archives/ Libraries Czech Department of Archives Administration CENDARI Data Sharing Agreement and Istituto Centrale per gli Archivi Data License The German Bundesarchiv DARIAH Memorandum of Understanding The National Archives of Estonia (Rahvusarhiiv) (MoU) Local Archives/ Libraries Letter of Collaboration Museo Storico Italiano della Guerra di Rovereto CENDARI Technical Checklist American JDC archives CENDARI Workflow Model 31 CONCLUSION JIRA 34 APPENDIX 16 BRINGING THE METADATA TOGETHER IN ONE REPOSITORY CENDARI/DARIAH Data Exchange Agree- ment 17 SOURCES AND NATURE OF CENDARI Frequently Asked Questions for CENDARIS ‘DATA SOUP’ Cultural Heritage Institutions Data Providers CENDARI Technical Checklist Categories of Data and Ingestion Options CENDARI Data Ingest Workflow Harvested (aggregated) data The data integration Dataspaces – the top level organization for CEN- DARI Data Integration of Dataspaces 4 5 The CENDARI White Book of Archives The CENDARI White Book of Archives EXECUTIVE SUMMARY Over the course of its four year project timeline, the CENDARI project has collected archi- val descriptions and metadata in various formats from a broad range of cultural heritage institutions. These data were drawn together in a single repository and are being stored there. The repository contains curated data which has been manually established by the CENDARI team as well as data acquired from small, ‘hidden’ archives in spreadsheet for- mat or from big aggregators with advanced data exchange tools in place. While the acquisition and curation of heterogeneous data in a single repository presents a technical challenge in itself, the ingestion of data into the CENDARI repository also opens up the possibility to process and index them through data extraction, entity recognition, semantic enhancement and other transformations. In this way the CENDARI project was able to act as a bridge between cultural heritage institutions and historical researchers, insofar as it drew together holdings from a broad range of institutions and enabled the browsing of this heterogeneous content within a single search space. This paper describes a broad range of ways in which the CENDARI project acquired data from cultural heritage institutions as well as the necessary technical background. In ex- emplifying diverse data creation or acquisition strategies, multiple formats and technical solutions, assets and drawbacks of a repository, this “White Book” aims at providing guid- ance and advice as well as best practices for archivists and cultural heritage institutions collaborating or planning to collaborate with infrastructure projects. 6 7 The CENDARI White Book of Archives The CENDARI White Book of Archives THE CENDARI WHITE BOOK OF ARCHIVES INTRODUCTION There is a wealth of historical material stored in Europe’s cultural heritage institutions. Thousands of records relevant for research are held within every one; the physical arte- facts preserved and organised so as to expose their provenance. Each record group can be consulted via analogue inventories and other finding aids, which are fully accessible within a given institution and often fully accessible online as well. The development of these vir- tual resources has occurred in concert with changes in historical methodology, where the pursuit of transnational topics and application of techniques such as data mining require and privilege digital sources that are fully exposed and widely available. In general, libraries and museums are at a much more advanced stage than archives in sharing and presenting their holdings to the broader public. These institutions have long since formed clusters to exchange data on the material they hold. More recently archives have started to apply similar methods and to build collaborations and networks in order to promote the development of their digital presence. Many institutions are still in the process of digitizing their analogue catalogues and inventories and are far from complet- ing this task. From the historian’s perspective, therefore, that landscape of provision can be quite uneven: while some institutions provide excellent digital access, others may display only some unstructured information in PDF or Word format on their websites. This variability in provision has its roots in the lack of resources faced by many cultural herit- age institutions, as well as in institutional cultures well-adapted to analogue processes and hierarchies, which may be different from their digital or virtual era equivalents. These restrictions can hold the institutions back from gaining the full benefit of the ‘digital turn’ in historical scholarship, and in particular from being able to fully participate in research infrastructure projects like CENDARI. CENDARI – the Collaborative European Digital Archival Research Infrastructure – began its work in 2012, charged with a mandate to ‘integrate digital archival resources for medieval and modern history’. In order to deliver on CENDARI’s case studies of World War I and Me- dieval Culture, the CENDARI team identified and contacted more than 250 cultural herit- age institutions whose collections were deemed pivotal enablers for research, and priority collections to expose in the CENDARI Virtual Research Environment (VRE). Apart from desk research, the CENDARI team contacted cultural heritage institutions to establish a basis for data exchange, i.e. collection descriptions already existing in digital formats and dis- played on individual websites of the institutions. The aim was to bring them together in one single repository – the CENDARI repository – thus facilitating the enhanced processing and enrichment of the data. At the same time, these contacting activities, including Skype conferences, phone calls, letters and on-site visits, were used to trace less visible archival information relevant to the CENDARI case studies, and to provide digital descriptions of collections where they didn’t yet exist. Because the cultural hertitage institutions involved are major research archives, muse- ums and libraries, they all had some digital presence or ongoing digitisation activities. This 8 9 The CENDARI White Book of Archives The CENDARI White Book of Archives presence did not guarantee the full display of information on all relevant collections and Altogether more than 1,200 archival institutions were added to CENDARI’s repository, us- holdings, nor – and more importantly – did the existence of digital finding aids and collec- ing an international standard format, the Encoded Archival Guide (EAG). tions mean that these resources could be accessed as ‘open data’ (that is, without restric- tions for their reuse) by a search engine from outside of the institution’s own closed data In a second step, more granular descriptions of specific relevant archival holdings were environment, e.g. via an application programming interface (API). The omnipresence of produced by CENDARI team members. Given the large number of relevant institutions such difficulties challenged the CENDARI team to develop a robust approach to data -ac identified and the breadth of their individual holdings, the project adopted a strategy to quisition and data management, satisfying the needs of both the institutions holding the prioritise 1) collections most relevant for the international research community and 2) data and the potential future users of CENDARI’s archival descriptions. collections most relevant for the Archival Research Guides (ARGs) to be produced by the project. In this way, the project captured both the key ‘backbone’ collections, which would It is the purpose of this document – the CENDARI White Book