Evaluating User Interest Profiles Using Ad Preference Managers
Total Page:16
File Type:pdf, Size:1020Kb
Quantity vs. Quality: Evaluating User Interest Profiles Using Ad Preference Managers Muhammad Ahmad Bashiry, Umar Farooqx, Maryam Shahidx, Muhammad Fareed Zaffarx, Christo Wilsony yNortheastern University, xLUMS (Pakistan) y{ahmad, cbw}@ccs.neu.edu, x{18100155, 18100048, fareed.zaffar}@lums.edu.pk Abstract—Widely reported privacy issues concerning major these profiles accurate? The extant literature presents a very online advertising platforms (e.g., Facebook) have heightened complete picture of who is tracking users’ online [7], [27], concerns among users about the data that is collected about [9], [86], [71], [70], as well as how tracking is implemented them. However, while we have a comprehensive understanding (e.g., fingerprinting) [37], [78], [5], [57], [60], [72], [58], [1], who collects data on users, as well as how tracking is implemented, [40], [28], but not necessarily the what. Controlled studies there is still a significant gap in our understanding: what have shown that advertising platforms like Google do indeed information do advertisers actually infer about users, and is this information accurate? draw inferences about users from tracking data [89], [20], [46], [47], but this does not address the broader question of what In this study, we leverage Ad Preference Managers (APMs) as platforms actually know about users in practice. a lens through which to address this gap. APMs are transparency tools offered by some advertising platforms that allow users to Answering this question is critical, as it directly speaks to see the interest profiles that are constructed about them. We the motivations of the advertising industry and its opposition recruited 220 participants to install an IRB approved browser to privacy-preserving regulations. Specifically, the industry extension that collected their interest profiles from four APMs claims that tracking and data aggregation are necessary to (Google, Facebook, Oracle BlueKai, and Neilsen eXelate), as well present users with relevant, targeted advertising. However, if as behavioral and survey data. We use this data to analyze the size user interest profiles are incorrect, this would suggest that and correctness of interest profiles, compare their composition across the four platforms, and investigate the origins of the data (1) online ads are being mistargeted and the money spent on underlying these profiles. them may be wasted, and (2) that greater privacy for users would have less of an impact on engagement with ads than is feared by industry. Further, surveys have found that users are I. INTRODUCTION deeply concerned when advertisers draw incorrect inferences If online advertising is the engine of the online economy, about them [24], a situation that investigative journalists have then data is the fuel for that engine. It is well known that anecdotally found to be true in practice [10], [48]. users’ behaviors are tracked as they browse the web [50], [27] In this study, we leverage Ad Preference Managers (APMs) and interact with smartphone apps [71], [70] to derive interest as a lens through which we answer these questions. APMs profiles that can be used to target behavioral advertising. Addi- are transparency tools offered by some advertising platforms tional sources of data like search queries, social media profiles (e.g., Google and Facebook) that allow users to see the interest and “likes”, and even offline data aggregated by brokers like profiles that advertisers have constructed about them. Most of Axciom [10], [48], are also leveraged by advertising platforms this data corresponds to user interests (e.g., “sports” or “com- to expand their profiles of users, all in an effort to sell more puters”), although some corresponds to user demographics and expensive ads that are more specifically targeted. attributes (e.g., “married” or “homeowner”). Widely reported privacy issues concerning major advertis- We recruited 220 participants to install an IRB approved ing platforms [19], [18] have heightened concerns among users browser extension that collected a variety of data (with user regarding their digital privacy. Several studies [55], [6], [84] consent). First, it gathered the full interest profile from four have found that users are uncomfortable with pervasive track- APMs: Google, Facebook, Oracle BlueKai, and Neilsen eXe- ing and the lack of transparency surrounding these practices. late. Second, it asked the participant to complete a survey on For example, a study by Pew found that 50% of online users demographics, general online activity, and privacy-conscious are concerned about the amount of information that is available behaviors (e.g., use of ad-blocking browser extensions). Third, about them online [69]. the extension showed participants a subset of their interests Despite widespread concerns about online privacy and at- from the APMs, and asked the participant if they were actually tention on the online advertising industry, there is one specific interested in these topics, as well as whether they recalled aspect of this ecosystem that remains opaque: what information seeing online ads related to these topics. Finally, the extension is actually contained in advertising interest profiles, and are collected the participant’s recent browsing history and history of Google Search queries. Network and Distributed Systems Security (NDSS) Symposium 2019 We use this mix of observational and qualitative data to 24-27 February 2019, San Diego, CA, USA ISBN 1-891562-55-X investigate several aspects of the interest profiles collected by https://dx.doi.org/10.14722/ndss.2019.23392 these four platforms, including: How many interests do the www.ndss-symposium.org APMs collect about users? Does each APM know the same interests about a given individual? and Are the interests (and ads targeted to them) actually relevant to users? Further, by place display ads within websites and mobile apps. Google is leveraging the last 100 days of historical data collected from the largest purveyor of keyword-based ads via Google Search our participants, we investigate what fraction of interests in the (i.e., AdWords) and video ads via YouTube. profiles could have been inferred from web tracking. We also Google’s advertising products draw data from a variety use regression models to examine whether privacy-conscious of sources. Studies have consistently shown that resources behaviors and use of privacy tools have any impact on the size from Google (e.g., Google Analytics and Tag Manager) are of interest profiles. embedded on >80% of domains across the web [31], [13], We make the following key contributions and observations [27], giving Google unprecedented visibility into users’ brows- throughout this study: ing behavior. Google collects usage data from Chrome and Android, the world’s most popular web browser and mobile • We present the first large-scale study of user interest operating system, respectively [87], [76]. Google’s trove of profiles that covers multiple platforms. user data includes real-time location (Android and Maps), text • We observe that interest profile sizes vary from zero to (Gmail, etc.), and voice transcripts (Assistant, etc.). thousands of interests per user, with Facebook having 1 the largest profiles by far. Because Facebook’s profiles Google’s APM has been available since 2011 [64]. As of are so large, they typically have 59–76% overlap with May 2018, this website allows a user to view lists of “topics” the other platforms; in contrast, Google, BlueKai, and that the user does and does not like. The lists are primarily eXelate’s profiles typically have ≤25% overlap for a populated based on inferences from users’ behavior, but users given user, i.e., different platforms have very different are able to add and remove topics from either list manually. “portraits” of users. Table I shows frequent and random interests that we • Participants were strongly interested in only 27% of observed in Google’s APM (the origins of our dataset are interests in their profiles. Further, participants reported described in § III). We see that Google’s inferred interests that they did not find ads targeted to low-relevance tend to be fairly broad, although some are more specific (e.g., interests to be useful. This finding suggests that many “Cricket” instead of “Sports”). As we will see, other APMs ads may be mistargeted, and highlights the need for present much more granular interests. a comprehensive study to evaluate the effectiveness of behavioral/targeted ads as compared to contex- Google’s policy states that they “do not show ads based tual/untargeted ads. on sensitive information or interests, such as those based on race, religion, sexual orientation, health, or sensitive financial • We find that recent browsing history only explains categories” [35]. However, two audit studies have cast doubt <9% of the interests in participants’ Facebook, on these claims, by demonstrating that Google did appear to BlueKai, and eXelate profiles (in the median case), infer sensitive attributes and allow advertisers to target them while recent browsing and search history can only for ads [89], [20]. More troubling, these inferred interests did explain 45% of participants’ Google profiles. Addi- not appear in Google’s APM. tionally, we find almost no significant correlations between privacy-conscious behaviors and interest pro- Facebook. As of May 2018, Facebook is the second largest file size. These findings suggest that non-tracking online advertiser by revenue [30]. Facebook primarily facili- data (data brokers) and other means of tracking like tates display and video advertising within its own products, browser fingerprinting and cross-device tracking may including its social network, Instagram, etc. Facebook closed be critical for building user interest profiles. It also their public web ad exchange in 2016, but still operates an suggests that alternate strategies and tools are needed exchange for in-app mobile advertising [65]. to protect users’ privacy. Facebook’s largest source of data are the personal profiles Outline. Our study is structured as follows: we begin that its users curate about themselves. This includes static by providing background on APMs in§II.