Wikipedia and Artificial Intelligence: an Evolving Synergy
Total Page:16
File Type:pdf, Size:1020Kb
Load more
Recommended publications
-
Cultural Anthropology Through the Lens of Wikipedia: Historical Leader Networks, Gender Bias, and News-Based Sentiment
Cultural Anthropology through the Lens of Wikipedia: Historical Leader Networks, Gender Bias, and News-based Sentiment Peter A. Gloor, Joao Marcos, Patrick M. de Boer, Hauke Fuehres, Wei Lo, Keiichi Nemoto [email protected] MIT Center for Collective Intelligence Abstract In this paper we study the differences in historical World View between Western and Eastern cultures, represented through the English, the Chinese, Japanese, and German Wikipedia. In particular, we analyze the historical networks of the World’s leaders since the beginning of written history, comparing them in the different Wikipedias and assessing cultural chauvinism. We also identify the most influential female leaders of all times in the English, German, Spanish, and Portuguese Wikipedia. As an additional lens into the soul of a culture we compare top terms, sentiment, emotionality, and complexity of the English, Portuguese, Spanish, and German Wikinews. 1 Introduction Over the last ten years the Web has become a mirror of the real world (Gloor et al. 2009). More recently, the Web has also begun to influence the real world: Societal events such as the Arab spring and the Chilean student unrest have drawn a large part of their impetus from the Internet and online social networks. In the meantime, Wikipedia has become one of the top ten Web sites1, occasionally beating daily newspapers in the actuality of most recent news. Be it the resignation of German national soccer team captain Philipp Lahm, or the downing of Malaysian Airlines flight 17 in the Ukraine by a guided missile, the corresponding Wikipedia page is updated as soon as the actual event happened (Becker 2012. -
Universality, Similarity, and Translation in the Wikipedia Inter-Language Link Network
In Search of the Ur-Wikipedia: Universality, Similarity, and Translation in the Wikipedia Inter-language Link Network Morten Warncke-Wang1, Anuradha Uduwage1, Zhenhua Dong2, John Riedl1 1GroupLens Research Dept. of Computer Science and Engineering 2Dept. of Information Technical Science University of Minnesota Nankai University Minneapolis, Minnesota Tianjin, China {morten,uduwage,riedl}@cs.umn.edu [email protected] ABSTRACT 1. INTRODUCTION Wikipedia has become one of the primary encyclopaedic in- The world: seven seas separating seven continents, seven formation repositories on the World Wide Web. It started billion people in 193 nations. The world's knowledge: 283 in 2001 with a single edition in the English language and has Wikipedias totalling more than 20 million articles. Some since expanded to more than 20 million articles in 283 lan- of the content that is contained within these Wikipedias is guages. Criss-crossing between the Wikipedias is an inter- probably shared between them; for instance it is likely that language link network, connecting the articles of one edition they will all have an article about Wikipedia itself. This of Wikipedia to another. We describe characteristics of ar- leads us to ask whether there exists some ur-Wikipedia, a ticles covered by nearly all Wikipedias and those covered by set of universal knowledge that any human encyclopaedia only a single language edition, we use the network to under- will contain, regardless of language, culture, etc? With such stand how we can judge the similarity between Wikipedias a large number of Wikipedia editions, what can we learn based on concept coverage, and we investigate the flow of about the knowledge in the ur-Wikipedia? translation between a selection of the larger Wikipedias. -
Christians and Jews in Muslim Societies
Arabic and its Alternatives Christians and Jews in Muslim Societies Editorial Board Phillip Ackerman-Lieberman (Vanderbilt University, Nashville, USA) Bernard Heyberger (EHESS, Paris, France) VOLUME 5 The titles published in this series are listed at brill.com/cjms Arabic and its Alternatives Religious Minorities and Their Languages in the Emerging Nation States of the Middle East (1920–1950) Edited by Heleen Murre-van den Berg Karène Sanchez Summerer Tijmen C. Baarda LEIDEN | BOSTON Cover illustration: Assyrian School of Mosul, 1920s–1930s; courtesy Dr. Robin Beth Shamuel, Iraq. This is an open access title distributed under the terms of the CC BY-NC 4.0 license, which permits any non-commercial use, distribution, and reproduction in any medium, provided no alterations are made and the original author(s) and source are credited. Further information and the complete license text can be found at https://creativecommons.org/licenses/by-nc/4.0/ The terms of the CC license apply only to the original material. The use of material from other sources (indicated by a reference) such as diagrams, illustrations, photos and text samples may require further permission from the respective copyright holder. Library of Congress Cataloging-in-Publication Data Names: Murre-van den Berg, H. L. (Hendrika Lena), 1964– illustrator. | Sanchez-Summerer, Karene, editor. | Baarda, Tijmen C., editor. Title: Arabic and its alternatives : religious minorities and their languages in the emerging nation states of the Middle East (1920–1950) / edited by Heleen Murre-van den Berg, Karène Sanchez, Tijmen C. Baarda. Description: Leiden ; Boston : Brill, 2020. | Series: Christians and Jews in Muslim societies, 2212–5523 ; vol. -
Modeling Popularity and Reliability of Sources in Multilingual Wikipedia
information Article Modeling Popularity and Reliability of Sources in Multilingual Wikipedia Włodzimierz Lewoniewski * , Krzysztof W˛ecel and Witold Abramowicz Department of Information Systems, Pozna´nUniversity of Economics and Business, 61-875 Pozna´n,Poland; [email protected] (K.W.); [email protected] (W.A.) * Correspondence: [email protected] Received: 31 March 2020; Accepted: 7 May 2020; Published: 13 May 2020 Abstract: One of the most important factors impacting quality of content in Wikipedia is presence of reliable sources. By following references, readers can verify facts or find more details about described topic. A Wikipedia article can be edited independently in any of over 300 languages, even by anonymous users, therefore information about the same topic may be inconsistent. This also applies to use of references in different language versions of a particular article, so the same statement can have different sources. In this paper we analyzed over 40 million articles from the 55 most developed language versions of Wikipedia to extract information about over 200 million references and find the most popular and reliable sources. We presented 10 models for the assessment of the popularity and reliability of the sources based on analysis of meta information about the references in Wikipedia articles, page views and authors of the articles. Using DBpedia and Wikidata we automatically identified the alignment of the sources to a specific domain. Additionally, we analyzed the changes of popularity and reliability in time and identified growth leaders in each of the considered months. The results can be used for quality improvements of the content in different languages versions of Wikipedia. -
Omnipedia: Bridging the Wikipedia Language
Omnipedia: Bridging the Wikipedia Language Gap Patti Bao*†, Brent Hecht†, Samuel Carton†, Mahmood Quaderi†, Michael Horn†§, Darren Gergle*† *Communication Studies, †Electrical Engineering & Computer Science, §Learning Sciences Northwestern University {patti,brent,sam.carton,quaderi}@u.northwestern.edu, {michael-horn,dgergle}@northwestern.edu ABSTRACT language edition contains its own cultural viewpoints on a We present Omnipedia, a system that allows Wikipedia large number of topics [7, 14, 15, 27]. On the other hand, readers to gain insight from up to 25 language editions of the language barrier serves to silo knowledge [2, 4, 33], Wikipedia simultaneously. Omnipedia highlights the slowing the transfer of less culturally imbued information similarities and differences that exist among Wikipedia between language editions and preventing Wikipedia’s 422 language editions, and makes salient information that is million monthly visitors [12] from accessing most of the unique to each language as well as that which is shared information on the site. more widely. We detail solutions to numerous front-end and algorithmic challenges inherent to providing users with In this paper, we present Omnipedia, a system that attempts a multilingual Wikipedia experience. These include to remedy this situation at a large scale. It reduces the silo visualizing content in a language-neutral way and aligning effect by providing users with structured access in their data in the face of diverse information organization native language to over 7.5 million concepts from up to 25 strategies. We present a study of Omnipedia that language editions of Wikipedia. At the same time, it characterizes how people interact with information using a highlights similarities and differences between each of the multilingual lens. -
Arxiv:2010.11856V3 [Cs.CL] 13 Apr 2021 Questions from Non-English Native Speakers to Rep- Information-Seeking Questions—Questions from Resent Real-World Applications
XOR QA: Cross-lingual Open-Retrieval Question Answering Akari Asaiº, Jungo Kasaiº, Jonathan H. Clark¶, Kenton Lee¶, Eunsol Choi¸, Hannaneh Hajishirziº¹ ºUniversity of Washington ¶Google Research ¸The University of Texas at Austin ¹Allen Institute for AI {akari, jkasai, hannaneh}@cs.washington.edu {jhclark, kentonl}@google.com, [email protected] Abstract ロン・ポールの学部時代の専攻は?[Japanese] (What did Ron Paul major in during undergraduate?) Multilingual question answering tasks typi- cally assume that answers exist in the same Multilingual document collections language as the question. Yet in prac- (Wikipedias) tice, many languages face both information ロン・ポール (ja.wikipedia) scarcity—where languages have few reference 高校卒業後はゲティスバーグ大学へ進学。 (After high school, he went to Gettysburg College.) articles—and information asymmetry—where questions reference concepts from other cul- Ron Paul (en.wikipedia) tures. This work extends open-retrieval ques- Paul went to Gettysburg College, where he was a member of the Lambda Chi Alpha fraternity. He tion answering to a cross-lingual setting en- graduated with a B.S. degree in Biology in 1957. abling questions from one language to be an- swered via answer content from another lan- 生物学 (Biology) guage. We construct a large-scale dataset built on 40K information-seeking questions Figure 1: Overview of XOR QA. Given a question in across 7 diverse non-English languages that Li, the model finds an answer in either English or Li TYDI QA could not find same-language an- Wikipedia and returns an answer in English or L . L swers for. Based on this dataset, we introduce i i is one of the 7 typologically diverse languages. -
(Anti-)Semitic” and “(Anti-) Zionist” in Modern Middle Eastern Discourse
On the use of the terms “(anti-)Semitic” and “(anti-) Zionist” in modern Middle Eastern discourse Lutz Edzard1 University of Oslo Abstract Just as political and cultural discourse in general, modern Middle Eastern discourse is at times characterized by a great deal of hostility, not only between different states or religious denominations, but also state-internally among various ethnic, political, or religious groups. This short article focuses on the use of the attributes “Semitic” and “Zionist,” as well as their negative counterparts “anti-Semitic” and “anti-Zionist,” respectively, in examples of both Arabic and Israeli critical to hostile discourse. The focus of the discussion will lie on how the original meanings of these terms, especially in their negated forms, tend to be distorted in engaged political and cultural discourse. Keywords: Semitic, Anti-Semitic, Zionist, lā-sāmī, ṣahyūnī, ʾanṭi-šemi, ṣiyyōn The term “Semitic” Brief overview of the term “Semitic” As is well known, the term “Semitic” derives from the name of one of the sons of Noah, Shem,2 and was suggested for the language family in question by the encyclopedic German historian and polymath August Ludwig von Schlözer in 1781. 3 In linguistics context, the term “Semitic” is generally speaking non-controversial. Together with Ancient Egyptian, as well as the language families Berber, Cushitic, Chadic, and possibly Omotic, the Semitic language family is part of the larger Afroasiatic macrofamily (formerly also referred to as “Hamito-Semitic”).4 This usage of the term “Semitic” must be kept apart from the usage of the term in the compound adjective “anti-Semitic,” a term only coined in 1879 in a pamphlet by the journalist Wilhelm Marr (if not already in 1860 by the bibliographer and Orientalist Moritz Steinschneider), referring to prejudices against or hatred of Jews. -
Explaining Cultural Borders on Wikipedia Through Multilingual Co-Editing Activity
Samoilenko et al. EPJ Data Science (2016)5:9 DOI 10.1140/epjds/s13688-016-0070-8 REGULAR ARTICLE OpenAccess Linguistic neighbourhoods: explaining cultural borders on Wikipedia through multilingual co-editing activity Anna Samoilenko1,2*, Fariba Karimi1,DanielEdler3, Jérôme Kunegis2 and Markus Strohmaier1,2 *Correspondence: [email protected] Abstract 1GESIS - Leibniz-Institute for the Social Sciences, 6-8 Unter In this paper, we study the network of global interconnections between language Sachsenhausen, Cologne, 50667, communities, based on shared co-editing interests of Wikipedia editors, and show Germany that although English is discussed as a potential lingua franca of the digital space, its 2University of Koblenz-Landau, Koblenz, Germany domination disappears in the network of co-editing similarities, and instead local Full list of author information is connections come to the forefront. Out of the hypotheses we explored, bilingualism, available at the end of the article linguistic similarity of languages, and shared religion provide the best explanations for the similarity of interests between cultural communities. Population attraction and geographical proximity are also significant, but much weaker factors bringing communities together. In addition, we present an approach that allows for extracting significant cultural borders from editing activity of Wikipedia users, and comparing a set of hypotheses about the social mechanisms generating these borders. Our study sheds light on how culture is reflected in the collective process -
Forgotten Palestinians
1 2 3 4 5 6 7 8 9 THE FORGOTTEN PALESTINIANS 10 1 2 3 4 5 6x 7 8 9 20 1 2 3 4 5 6 7 8 9 30 1 2 3 4 5 36x 1 2 3 4 5 6 7 8 9 10 1 2 3 4 5 6 7 8 9 20 1 2 3 4 5 6 7 8 9 30 1 2 3 4 5 36x 1 2 3 4 5 THE FORGOTTEN 6 PALESTINIANS 7 8 A History of the Palestinians in Israel 9 10 1 2 3 Ilan Pappé 4 5 6x 7 8 9 20 1 2 3 4 5 6 7 8 9 30 1 2 3 4 YALE UNIVERSITY PRESS 5 NEW HAVEN AND LONDON 36x 1 In memory of the thirteen Palestinian citizens who were shot dead by the 2 Israeli police in October 2000 3 4 5 6 7 8 9 10 1 2 3 4 5 Copyright © 2011 Ilan Pappé 6 The right of Ilan Pappé to be identified as author of this work has been asserted by 7 him in accordance with the Copyright, Designs and Patents Act 1988. 8 All rights reserved. This book may not be reproduced in whole or in part, in any form (beyond that copying permitted by Sections 107 and 108 of the U.S. Copyright 9 Law and except by reviewers for the public press) without written permission from 20 the publishers. 1 For information about this and other Yale University Press publications, 2 please contact: U.S. -
Classification of Wikipedia Articles Using BERT in Combination with Metadata
Daniel Garber¨ Classification of Wikipedia Articles using BERT in Combination with Metadata Bachelor’s Thesis to achieve the university degree of Bachelor of Science Bachelor’s degree programme: Soware Engineering and Management submied to Graz University of Technology Supervisor Dipl.-Ing. Dr.techn. Roman Kern Institute for Interactive Systems and Data Science Head: Univ.-Prof. Dipl.-Inf. Dr. Stefanie Lindstaedt Graz, August 2020 Affidavit I declare that I have authored this thesis independently, that I have not used other than the declared sources/resources, and that I have explicitly indicated all material which has been quoted either literally or by content from the sources used. The text document uploaded to tucnazonline is identical to the present master's thesis. Signature Abstract Wikipedia is the biggest online encyclopedia and it is continually growing. As its complexity increases, the task of assigning the appropriate categories to articles becomes more dicult for authors. In this work we used machine learning to auto- matically classify Wikipedia articles from specic categories. e classication was done using a variation of text and metadata features, including the revision history of the articles. e backbone of our classication model was a BERT model that was modied to be combined with metadata. We conducted two binary classication experiments and in each experiment compared various feature combinations. In the rst experiment we used articles from the category ”Emerging technologies” and ”Biotechnology”, where the best feature combination achieved an F1 score of 91.02%. For the second experiment the ”Biotechnology” articles are exchanged with random Wikipedia articles. Here the best feature combination achieved an F1 score of 97.81%. -
Critical Point of View: a Wikipedia Reader
w ikipedia pedai p edia p Wiki CRITICAL POINT OF VIEW A Wikipedia Reader 2 CRITICAL POINT OF VIEW A Wikipedia Reader CRITICAL POINT OF VIEW 3 Critical Point of View: A Wikipedia Reader Editors: Geert Lovink and Nathaniel Tkacz Editorial Assistance: Ivy Roberts, Morgan Currie Copy-Editing: Cielo Lutino CRITICAL Design: Katja van Stiphout Cover Image: Ayumi Higuchi POINT OF VIEW Printer: Ten Klei Groep, Amsterdam Publisher: Institute of Network Cultures, Amsterdam 2011 A Wikipedia ISBN: 978-90-78146-13-1 Reader EDITED BY Contact GEERT LOVINK AND Institute of Network Cultures NATHANIEL TKACZ phone: +3120 5951866 INC READER #7 fax: +3120 5951840 email: [email protected] web: http://www.networkcultures.org Order a copy of this book by sending an email to: [email protected] A pdf of this publication can be downloaded freely at: http://www.networkcultures.org/publications Join the Critical Point of View mailing list at: http://www.listcultures.org Supported by: The School for Communication and Design at the Amsterdam University of Applied Sciences (Hogeschool van Amsterdam DMCI), the Centre for Internet and Society (CIS) in Bangalore and the Kusuma Trust. Thanks to Johanna Niesyto (University of Siegen), Nishant Shah and Sunil Abraham (CIS Bangalore) Sabine Niederer and Margreet Riphagen (INC Amsterdam) for their valuable input and editorial support. Thanks to Foundation Democracy and Media, Mondriaan Foundation and the Public Library Amsterdam (Openbare Bibliotheek Amsterdam) for supporting the CPOV events in Bangalore, Amsterdam and Leipzig. (http://networkcultures.org/wpmu/cpov/) Special thanks to all the authors for their contributions and to Cielo Lutino, Morgan Currie and Ivy Roberts for their careful copy-editing. -
Art, History and the Historiography of Judaism in Roman Antiquity Brill Reference Library of Judaism
Art, History and the Historiography of Judaism in Roman Antiquity Brill Reference Library of Judaism Editors Alan J. Avery-Peck (College of the Holy Cross) William Scott Green (University of Rochester) Editorial Board David Aaron (Hebrew Union College-Jewish Institute of Religion, Cincinnati) Herbert Basser (Queen’s University) Bruce D. Chilton (Bard College) José Faur (Netanya College) Neil Gillman (Jewish Theological Seminary of America) Mayer I. Gruber (Ben-Gurion University of the Negev) Ithamar Gruenweld (Tel Aviv University) Maurice-Ruben Hayoun (University of Strasbourg and Hochschule fuer Juedische Studien, Heidelberg) Arkady Kovelman (Moscow State University) David Kraemer (Jewish Theological Seminary of America) Baruch A. Levine (New York University) Alan Nadler (Drew University) Jacob Neusner (Bard College) Maren Niehoff (Hebrew University of Jerusalem) Gary G. Porton (University of Illinois) Aviezer Ravitzky (Hebrew University of Jerusalem) Dov Schwartz (Bar Ilan University) Günter Stemberger (University of Vienna) Michael E. Stone (Hebrew University of Jerusalem) Elliot Wolfson (New York University) VOLUME 34 The titles published in this series are listed at brill.com/brlj Brill Reference Library Art, History and the Historiography of Judaism of Judaism in Roman Antiquity Editors By Alan J. Avery-Peck (College of the Holy Cross) Steven Fine William Scott Green (University of Rochester) Editorial Board David Aaron (Hebrew Union College-Jewish Institute of Religion, Cincinnati) Herbert Basser (Queen’s University) Bruce D. Chilton (Bard College) José Faur (Netanya College) Neil Gillman (Jewish Theological Seminary of America) Mayer I. Gruber (Ben-Gurion University of the Negev) Ithamar Gruenweld (Tel Aviv University) Maurice-Ruben Hayoun (University of Strasbourg and Hochschule fuer Juedische Studien, Heidelberg) Arkady Kovelman (Moscow State University) David Kraemer (Jewish Theological Seminary of America) Baruch A.