Semantic Web 0 (0) 1 1 IOS Press 1 1 2 2 3 3 4 WarSampo Knowledge Graph: Finland in the 4 5 5 6 Second World War as Linked Open Data 6 7 7 8 Mikko Koho a,b,*, Esko Ikkala a , Petri Leskinen a , Minna Tamper a,b , Jouni Tuominen a,b , and 8 9 Eero Hyvönen a,b 9 10 10 a Semantic Computing Research Group (SeCo), Aalto University, Department of Computer Science, Finland 11 11 E-mail: firstname.lastname@aalto.fi 12 12 b HELDIG – Helsinki Centre for Digital Humanities, University of Helsinki, Finland 13 13 E-mail: firstname.lastname@helsinki.fi 14 14 15 15 Editor: Christoph Schlieder, University of Bamberg, Germany 16 16 Solicited reviews: Laura Pandolfo, University of Sassari, Sassari, Italy; Günther Görz, AG Digital Humantities, 17 Friedrich-Alexander-University, Erlangen-nuremberg, Germany; One anonymous reviewer 17 18 18 19 19 20 20 21 21 22 22 Abstract. The Second World War (WW2) is arguably the most devastating catastrophe of human history, a topic of great interest 23 to not only researchers but the general public. However, data about the Second World War is heterogeneous and distributed in 23 24 various organizations and countries making it hard to utilize. In order to create aggregated global views of the war, a shared 24 25 ontology and data infrastructure is needed to harmonize information in various data silos. This makes it possible to share data 25 26 between publishers and application developers, to support data analysis in Digital Humanities research, and to develop data- 26 27 driven intelligent applications. As a first step towards these goals, this article presents the WarSampo knowledge graph (KG), a 27 28 shared semantic infrastructure, and a Linked Open Data (LOD) service for publishing data about WW2, with a focus on Finnish 28 29 military history. The shared semantic infrastructure is based on the idea of representing war as a spatio-temporal sequence of 29 30 events that soldiers, military units, and other actors participate in. The used metadata schema is an extension of CIDOC CRM, 30 31 supplemented by various military history domain ontologies. With an infrastructure containing shared ontologies, maintaining 31 the interlinked data brings upon new challenges, as one change in an ontology can propagate across several datasets that use it. To 32 32 support sustainability, a repeatable automatic data transformation and linking pipeline has been created for rebuilding the whole 33 33 WarSampo KG from the individual source datasets. The WarSampo KG is hosted on a data service based on W3C Semantic 34 Web standards and best practices, including content negotiation, SPARQL API, download, automatic documentation, and other 34 35 services supporting the reuse of the data. The WarSampo KG, a part of the international LOD Cloud and totalling ca. 14 million 35 36 triples, is in use in nine end-user application views of the WarSampo portal, which has had over 690 000 end users since its 36 37 opening in 2015. 37 38 38 39 Keywords: Linked Open Data, Semantic Web, Military History, World War II, Finland, Cultural Heritage, Digital Humanities 39 40 40 41 41 42 42 43 1. Introduction: Military History as Linked Data zations, making it hard to find and use. Furthermore, 43 44 the information is seldom published as data for reuse 44 45 Plenty of information about WW2 is published ev- in computational analyses and applications. Gather- 45 46 ery year in books, articles, news, web sites and ser- ing, extracting, and harmonizing information about the 46 47 vices, documentaries, and films for humans to con- war is needed in order to create comprehensive global 47 48 sume. This information is scattered in various mili- views of the war and history but this is not a simple 48 49 tary, governmental, cultural heritage, and other organi- task. This applies also to microhistory: for example, 49 50 finding out the details of what happened to a perished 50 51 *Corresponding author. E-mail: firstname.lastname@aalto.fi. relative during the war can be quite tedious, involving 51 1570-0844/0-1900/$35.00 c 0 – IOS Press and the authors. All rights reserved 2 M. Koho et al. / WarSampo Knowledge Graph: Finland in WW2 as Linked Open Data 1 studying and aggregating data about him/her from sev- have used the CIDOC Conceptual Reference Model 1 2 eral registries and data sources. Without harmonized, (CRM)5 [13]. 2 3 clean data, the data analysis of large military historical Several projects have published linked data about 3 4 datasets, such as death records, would be difficult in the World War I on the web, such as Europeana Collec- 4 5 Digital Humanities Research [1, 2]. Combining infor- tions 1914–19186, 1914–1918 Online7, WW1 Discov- 5 6 mation from various sources facilitates answering the ery8, CENDARI9 [14], Muninn10, and WW1LOD [15]. 6 7 complex societal research questions of “new military There are also a few works that have used the Linked 7 8 history” scholars [3, 4]. Data approach to WW2, such as [16–18] and a LOD 8 9 WarSampo Initiative and Project Series. The system on WW2 holocaust victims [19]. 9 10 goal of the WarSampo – Finnish Second World War Our own previous research on WarSampo first pre- 10 11 on the Semantic Web initiative1 is to study and show sented the vision and overview of the system especially 11 12 how Linked Data [5] (LD) can help in solving tasks from the use case and end-user application perspec- 12 13 like these [6]. The initiative collects military histori- tives [6, 20]. In [21] data integration was concerned 13 14 cal data related to Finland in the Second World War from the named entity linking (NEL) point of view. 14 15 (WW2). The data is published as Linked Open Data The maintenance problem of the interlinked dataset 15 16 (LOD) in an open SPARQL endpoint on top of which has been explored in [22]. Work on creating and using 16 17 the WarSampo portal2 has been created, including nine individual parts of the KG has been presented in sev- 17 18 application perspectives to the data. The portal, tar- eral previous publications [7, 8, 23–26]. This dataset 18 19 geted to both researchers and the public at large was description complements our other publications about 19 20 opened in 2015. The WarSampo data service and por- WarSampo by presenting in detail the KG, including 20 21 tal were awarded with the LODLAM Challenge Open the process of maintaining the data. 21 22 Data Prize in 2017 in Venice. The data forms an inte- This article is organized as follows. The next Sec- 22 23 grated interlinked 5-star LOD publication, and is part tion presents the source datasets. Section 3 discusses 23 24 of the global LOD Cloud3. how the information in the source datasets was har- 24 25 The WarSampo knowledge graph (KG) was pub- monized and presents the event-based data model. The 25 26 lished initially in 2015. The KG was first used by data transformation process is presented in Section 4. 26 27 seven different application perspectives in the War- An analysis of the data quality is given in Section 5. 27 28 Sampo portal, via only the SPARQL API [6]. The idea The stability and usefulness of the data are discussed in 28 29 was to show that anyone could easily use the data dy- Sections 6 and 7, respectively. Conclusion is provided 29 30 namically on the client side. In 2017, by the centen- in Section 8. 30 31 nial of Finnish independence, a new eighth applica- 31 32 tion perspective of war cemetery data and related pho- 32 33 tographs4 was released [7], a further demonstration of 2. Source Datasets 33 34 this idea. Finally, a ninth application based on a dataset 34 Table 1 lists the heterogeneous source datasets of 35 of 4200 prisoners of war was aligned with the War- 35 WarSampo. The data comes from several Finnish or- 36 Sampo KG and was released [8] in November 2019. 36 ganizations, such as the National Archives of Fin- 37 Related Work. The problem of combining and us- 37 land, the Finnish Defence Forces, and the National 38 ing heterogeneous cultural heritage datasets is a com- 38 Land Survey of Finland. Some source datasets have 39 mon problem in using Linked Data for Digital Hu- 39 been created as part of the WarSampo project and 40 manities [9, 10] and in Digital History [11]. His- 40 related research. The source datasets are in different 41 torical knowledge contextualization and visualization 41 formats, e.g., spreadsheets, text, web pages, images, 42 with experiences from the VICODI project are repre- 42 application programming interfaces (API), Extensible 43 sented in [12], which also discusses general problems 43 Markup Language (XML) documents, Portable Doc- 44 faced when modelling history with ontologies. Sev- 44 45 eral humanities and cultural heritage related projects 45 46 5A list of CIDOC CRM use cases can be found at: http://www. 46 47 cidoc-crm.org/useCasesPage. 47 1 6 48 The initiative and publications are presented in the initiative http://www.europeana-collections-1914-1918.eu 48 homepage: https://seco.cs.aalto.fi/projects/sotasampo/en/. 7http://www.1914-1918-online.net 49 49 2http://sotasampo.fi/en 8http://ww1.discovery.ac.uk 50 3http://linkeddata.org 9http://www.cendari.eu 50 51 4https://seco.cs.aalto.fi/projects/sotasampo/hautausmaat/ 10http://blog.muninn-project.org 51 M. Koho et al. / WarSampo Knowledge Graph: Finland in WW2 as Linked Open Data 3 1 ument Format (PDF) documents, and Resource De- the heterogeneous source data for a unified represen- 1 2 scription Framework (RDF) graphs.
Details
-
File Typepdf
-
Upload Time-
-
Content LanguagesEnglish
-
Upload UserAnonymous/Not logged-in
-
File Pages14 Page
-
File Size-