A Wikia Census: Motives, Tools and Insights
Total Page:16
File Type:pdf, Size:1020Kb
A Wikia census: motives, tools and insights Guillermo Jimenez-Diaz Abel Serrano Javier Arroyo Facultad de Informática. Universidad Facultad de Informática. Universidad Facultad de Informática. Universidad Complutense de Madrid Complutense de Madrid Complutense de Madrid Madrid, Spain Madrid, Spain Madrid, Spain [email protected] [email protected] [email protected] ABSTRACT have noted [2]. For example, the growth stages experimented by Understanding the Wikisphere phenomenon is undoubtedly of Wikipedia may be different to those from other wikis; or the orga- great interest. Most of the studies focus in Wikipedia and its gener- nization patterns in Wikipedia projects may not emerge in other alization to other wikis requires an enormous amount of work in projects. Furthermore, the study of diverse wikis can shed light terms of selecting and retrieving the data. To facilitate the analysis on project sustainability, stagnation, or abandonment, which are of other wikis we developed a set of tools to collect and create a cen- critical phenomenon for peer production that cannot be studied in sus of Wikia, one of the largest and most diverse repository of wikis, a mature successful project such as Wikipedia. which hosts more than 300,000 wikis. In this work, we carry out a Hence, research on knowledge peer production using wikis preliminary quantitative analysis of the census, emphasizing on the should be based on the diversity of the wikis in the Internet, the so- differences between active and inactive wikis. Additionally, wepro- called wikisphere. That would require to change the object of study vide the wiki research community with the census and the scripts from a single wiki or a few of them, to a numerous sample of them. employed to retrieve the data, facilitating others to reproduce or The quantitative analysis of wikis may help to better understand reuse it. these socio-technical systems in its diversity. Some notable exam- ples of research on the wikisphere include [2, 5, 7–9]. However, this CCS CONCEPTS kind of studies is still scarce and hence much remains to be known about knowledge production in wikis from a broad scope. • Information systems → Wikis; • Human-centered comput- One of the causes of this scarcity may well be the lack of quan- ing → Wikis; Collaborative and social computing design and eval- titative wiki data, as noted by [8]. Each single research, such as uation methods; those cited above, implied a significant amount of work in terms KEYWORDS of selecting and retrieving the data. However, since data are not publicly available, they cannot be reused by other researchers. Thus, Knowledge P2P production, wikis, Wikia, census, wikisphere, on- as of today, there is no curated corpus of wikis available to prompt line communities, collaborative work. quantitative research in online collaborative knowledge production. ACM Reference Format: Not even the tools used to retrieve it are publicly available, although Guillermo Jimenez-Diaz, Abel Serrano, and Javier Arroyo. 2018. A Wikia cen- they may be upon request to their respective authors. sus: motives, tools and insights. In OpenSym ’18: The 14th International Sym- Another cause may be the fact that the characteristics and di- posium on Open Collaboration, August 22–24, 2018, Paris, France. ACM, New mensions of the population under study, the so-called wikisphere, York, NY, USA, 6 pages. https://doi.org/10.1145/3233391.3233526 are not known and probably cannot be known but approximately: how many wikis are there? how big are they? how many people 1 MOTIVATION AND AIMS collaborate in them? Wikis have enabled the radical change in terms of peer-production The task is daunting, but efforts even if partial can still be of great of knowledge and open collaboration that has taken place in this value as they serve to better define some regions of the wikisphere. century. From all the peer-produced content projects, Wikipedia It could help to carry out scientific studies on such regions, making is the flagship example as it is the world’s leading source ofweb possible to define a sampling process to retrieve a sample of wikis reference information [3]. As a result, it has drawn attention from from the whole population. One example is the ambitious resource researchers all around the world and it is being studied from both developed by the collective behind s23.org. This website collected qualitative and quantitative perspectives [1, 4]. metrics of wikis from MediaWiki-based repositories totaling over However, findings from the most outstanding wiki project in 38,000 wikis. A brief demographic analysis on these data can be the world may not generalize to other wikis, as some researchers found in [5]. However, the data on the website have not been up- dated since 2012 or 2014 (depending on the repositories) and the Permission to make digital or hard copies of all or part of this work for personal or tools to collect the data are not available in public repositories to classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation the best of our knowledge. on the first page. Copyrights for components of this work owned by others than the This work aims to help stimulating the study of the wikisphere author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or and to make a contribution useful for the wiki research community. republish, to post on servers or to redistribute to lists, requires prior specific permission 1 and/or a fee. Request permissions from [email protected]. More precisely, we present a census of the wikis from Wikia and OpenSym ’18, August 22–24, 2018, Paris, France make available the tools used to retrieve it. © 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-5936-8/18/08...$15.00 https://doi.org/10.1145/3233391.3233526 1Also known as Fandom: http://www.wikia.com/ OpenSym ’18, August 22–24, 2018, Paris, France G. Jimenez-Diaz et al. Our census focuses on Wikia, so it does not represent the whole research on wikis and on open-knowledge production in general wikisphere, but an important part of it, as Wikia is the largest and to contribute to the reproducibility of its experiments. repository of wikis to date. In Wikia the content of the wikis is open-licensed and free and its creation is also free of charge. The 2 THE CENSUS OF WIKIA downside for users may be the advertisement banners that appear Besides each name and web domain (URL), the census that we in the wikis, which can make some people or communities to look present here includes the following information for each wiki: for other wiki hosting service to create their wikis. However, it is • Number of the revisions made in all the pages contained in undoubtedly a paramount example of open knowledge creation the wiki (Edits). peer communities • Number of content pages in the wiki (Articles). Interestingly, most wikis from Wikia are spontaneously created • Number of all pages in the wiki, including articles, talk pages, by users, unlike wikis from Wikimedia Foundation or others hosted redirects, etc. (Pages). by public or private organizations. As a result, the topics of Wikia • Wiki Language (Lang). wikis vary enormously, mainly addressing popular culture and • Category of the wiki contents, according to a fixed set of fan culture (video-games, films, TV shows, books, comics, etc.), categories defined by Wikia (Hub) and by a user-defined but covering a diversity of other topics as well, such as hobbies category (Topic). (food, fashion, technology, genealogy, collaborative fiction creation, • Number of users who have performed an action in the last etc.), collaborative resources (an encyclopedia of psychology, a 30 days (ActiveUsers). protein taxonomy for biologists, a genealogy, a song lyrics archive, • WAM Score, a popularity metric computed by Wikia as a etc.), or related with activism (about LGTB, sustainability, political combination of traffic, engagement and growth (WAM)2. orientations, etc). The diversity in terms of languages represented is • Number of human registered users, categorized according also wide. Finally, in Wikia there are wikis with very different ages, to the number of edits in the last 30 days (Users_N). number of articles and community sizes. In this sense, it is a valuable • Estimated date of creation of the wiki (BirthDate). resource to study topics such as maturity, growth, stagnation and The census is publicly available at https://www.kaggle.com/abeserra/ abandonment, all of them relevant issues for peer production and wikia-census. community health understanding. Thus, a census of Wikia will offer a detailed description of an im- 2.1 Description of the process portant and diverse region of the wikisphere that, otherwise, is only known by guesses and approximations. For example, one may guess The process of collecting the data census involves different stages Wikia hosts thousands of wikis mostly in English and mostly about and retrieving data from different sources. In this section, we will fan content (movies, TV shows, video-games, comics...) with some briefly describe the main aspects about the process. For amore notable examples such as the Harry Potter or Star Wars wikis. How- detailed account, we refer to the appendix in this article and to the 3 ever, such guesses, even if correct, are not empirically grounded code repository publicly available in Github . and tell us nothing about the rest of Wikia. In this work, we carry The names and URLs of all the wikis available in Wikia can be 4 out a brief analysis of the retrieved census, so we are able to offer an found in the sitemap page of Wikia .