Analyzing Reading Behavior by Blog Mining

Analyzing Reading Behavior by Blog Mining Tadanobu Furukawa Yutaka Matsuo Ikki Ohmukai Koki Uchiyama Mitsuru Ishizuka AIST National Institute of Informatics Hotto Link Inc. University of Tokyo 1-18-13 Sotokanda 2-1-2 Hitotsubashi 1-6-1 Otemachi 7-3-1 Hongo, Bunkyo-ku Chiyoda-ku, Tokyo Chiyoda-ku, Tokyo Chiyoda-ku, Tokyo Tokyo 135-8656, Japan 101-0021, Japan 101-8430, Japan 100-0004, Japan Abstract Adamic & Glance 2005). Among several studies that analyzed social networks on the blogosphere, some have used This paper presents a study of the various aspects of citation (mention of urls in the posting) to show relations blog reading behavior. The analyzed data are obtained among blogs; others have used blogrolls or trackbacks as from a Japanese weblog hosting service, Doblog. Four kinds of social networks are generated and analyzed: ci- evidence of relations. At least one study has surveyed com- tation, comment, trackback, and blogroll networks. In ment relations among bloggers (Lento et al. 2006). addition, the user log data are used to identify reader- Another important relation exists among bloggers: read- ship relations among bloggers. After analysis of more ership relations. Readership relations are not observable than 50,000 users for about two years, we reveal some through publicly available data, but they are an important interactions between social relations and readership re- source of information because bloggers read other blogs and lations. We first show that bloggers read other weblogs write their own blogs. The importance of readership rela- on a regular basis (50% of weblogs that are read at least tions is described in the literature (such as (Efimova, Hen- three times are read every five times a user logs in). We drick, & Anjewierden 2005; Nardi, Schiano, & Gumbrecht call this relation a regular reading relation (RR rela- 2004)). A recent study has analyzed the structural properties tion). Then, prediction of RR relations is done using features from the four kinds of social networks. Lastly, of weblog readership networks (Marlow 2006). information diffusion on RR relations is analyzed and Although those studies reveal various interesting findings characterized. Results of this study show that the blogs related to the social networks of weblogs, no comprehen- in RR relations have an important role in bloggers’ ac- sive study of them has examined all social and behavioral tivities. We find the features which have a correlation relations simultaneously. A comparison among different re- with RR relations. lations provides the general overview of each relation and its associated pattern of interaction. For this study, we analyze multiple social networks of weblogs: citation, com- Introduction ment, trackback, and blogroll networks. Subsequently, the Web logs (blogs, or weblogs) constitute a prominent social user log data (which include which blogs a user browses medium on the internet that enables users to publish indi- and when) are used to identify readership relations among vidual experiences and opinions easily. Analyzing these bloggers. Therefore, five types of social and behavioral net- data helps to understand the user’s behavioral pattern on the works are analyzed in this paper. We use the database of a blogosphere, and it supports the development of web ser- blog-hosting service in Japan called Doblog1. vices such as recommendation of blogs. Numerous studies Analyses of 1.5 million entries made by more than 50,000 have examined weblogs, especially addressing their social users for about two years reveal interesting interactions in- aspects: Bloggers read other blogs and leave comments and volving social relations and readerships. This paper de- send trackbacks as they update their own blogs. Users might scribes salient aspects of the following analysis. mention other blogs in their postings, and express their sug- • We first illustrate the four kinds of social networks and gested contacts in blogrolls (a sidebar within a particular characterize them. blog listing the other blogs the blogger frequents). These activities present an interwoven record of multiple relation- • Bloggers read some blogs on a regular basis. We can dis- ships among blogs (sometimes called a multiplex graph in cern these behaviors quantitatively: 50% of weblogs that sociology), which is an interesting source of information to are read at least three times are read every five times a characterize user behavior, community structure, and infor- user logs in. We call this relation a regular reading rela- mation diffusion. tion (RR relation). To date, various studies have specifically examined so- • Link prediction of RR relations is done using features cial network aspects of weblog authors (Marlow 2004; from the four kinds of social networks. Some attributes Copyright c 2007, Association for the Advancement of Artificial 1Doblog (http://www.doblog.com/), provided by NTT Intelligence (www.aaai.org). All rights reserved. Data Corp. and Hotto Link, Inc. 1353 (such as graph distance) and some networks (blogrolls Kraut 2002). Computer-mediated communication (in par- and citations) are demonstrably useful for predicting RR ticular, e-mail) is less valuable for building and sustaining relations. close social relationships than face-to-face contact and tele- • Information diffusion through RR relations is analyzed phone conversations. R. Kumar et al. investigates profiles and characterized. Some information is likely to be of more than one million livejournal.com bloggers in 2004, conveyed through RR relations. Generally, information and analyzes users’ demographic and geographic character- propagates in a shorter time and with higher probability istics (Kumar et al. 2004). More recently, Ali-Hasan and through RR relations than through non-RR relations. L. Adamic find interesting characteristics of bloggers’ online and real-life relationships (ALi-Hasan & Adamic 2007). Our findings provide an overview of social relations and They investigate three blog communities using an online reading behavior. These results support those of existing survey, which reveals that few blogging interactions reflect studies of social network analyses of the blogosphere. close offline relationships; furthermore, many online rela- This paper is organized as follows: first, we describe re- tionships were formed through blogging. lated studies. Next, we explains the definition of social Nardi et al. conducted audiotaped ethnographic inter- relations among weblogs, and also define the RR relation views with 23 bloggers, with analysis of their blog posts and characterize it with social relations. We then produce a (Nardi, Schiano, & Gumbrecht 2004). The motivations model to infer the existence of RR relations as a link predic- of blogging are enumerated. They provide good insights tion problem, and analyze information diffusion through RR that support the background of our research: bloggers write relations and show the effect of RR relations. Finally, after a blogs to (1) update others on activities and whereabouts, (2) discussion of analytical limitations, we conclude the paper. express opinions to influence others, (3) seek others’ opinions and feedback, (4) “think by writing”, and (5) release Related Works emotional tension. Nardi et al. remark that Many studies have specifically undertaken analysis of the Our research leads us to speculate that blogging is as blogosphere as a social medium: trend detection, network much about reading as writing, and as much about lis- analysis, user profiling, and splog (spam blog) detection. We tening as talking. We specifically examined the pro- introduce several works that are closely related to ours. duction of blogs, but future research will address blog Several studies have analyzed social networks that exist readers and to assess the relations between blog writers in the blogosphere: L. Adamic and N. Glance study the link and blog readers precisely. patterns (citations and blogrolls) and discussion topics of po- In the following section, we investigate relations between litical bloggers (Adamic & Glance 2005). This study de- blog writers and readers from a social network perspective. tected differences in the behaviors of politically liberal and conservative blogs, with conservative blogs linking to each Four Social Networks among Weblogs other more frequently. Lento et al. conduct data analyses re- garding the Wallop system and compares users who remain In this section, we define four kinds of relations among blogs active to those who do not (Lento et al. 2006). Similarly and depict social networks defined by these relations. to our work, the use of data from the hosting service enables We consider four types of relations between two blogs: them to detect social ties such as comment relations and invi- Citation We define that there is a citation relation from A tation relations, which are usually impossible to obtain. G. to B if an entry of blog A includes a hyperlink to blog B. Mishne and N. Glance analyze blog comments (Mishne & Blogroll We assert a blogroll relation from A to B if a Glance 2006). Those studies extract some relations among blogroll (a list of weblogs in the front page) of A includes blogs and produce a social network for analysis. blog B. Social networks are used for several applications such as Comment A comment relation from A to B pertains if the blog/entry ranking and community detection: E. Adar et al. blogger of blog A comments on blog B. proposes a ranking algorithm called iRank, which is based on implicit routes of information transmission as well as Trackback A trackback relation from A to B exists if an explicit links (Adar et al. 2004). For community detec- entry of blog B contains a back-reference by the trackback tion, Y. Lin et al. seek interesting aspects of social relations function to blog A. (Lin et al. 2006). They develop a computational model for These are mentioned in (Marlow 2004) and other literature. mutual awareness that incorporates specific action types in- We call these four relations social relations because the re- cluding commenting and changing blogrolls. The mutual lations are publicly observable and therefore involve some awareness feature is used for community extraction.

Analyzing Reading Behavior by Blog Mining

Uncovering Social Network Sybils in the Wild

Svms for the Blogosphere: Blog Identification and Splog Detection

By Nilesh Bansal a Thesis Submitted in Conformity with the Requirements

Blogosphere: Research Issues, Tools, and Applications

Web Spam Taxonomy

Robust Detection of Comment Spam Using Entropy Rate

Spam in Blogs and Social Media

Advances in Online Learning-Based Spam Filtering

2 Uncovering Social Network Sybils in the Wild

The SAGE Handbook of Web History by Niels Brügger, Ian

Social Software: Fun and Games, Or Business Tools?

D2.4 Weblog Spider Prototype and Associated Methodology