Combining Multiple Metadata Types in Movies Recommendation Using Ensemble Algorithms

Total Page:16

File Type:pdf, Size:1020Kb

Combining Multiple Metadata Types in Movies Recommendation Using Ensemble Algorithms Universidade de São Paulo Biblioteca Digital da Produção Intelectual - BDPI Departamento de Ciências de Computação - ICMC/SCC Comunicações em Eventos - ICMC/SCC 2014-11 Combining multiple metadata types in movies recommendation using ensemble algorithms Brazilian Symposium on Multimedia and the Web, 20th, 2014, João Pessoa. http://www.producao.usp.br/handle/BDPI/48641 Downloaded from: Biblioteca Digital da Produção Intelectual - BDPI, Universidade de São Paulo Combining Multiple Metadata Types in Movies Recommendation Using Ensemble Algorithms Bruno Cabral Renato D. Beltrão Marcelo G. Manzato Department of Computer Mathematics and Computing Mathematics and Computing Science Institute - University of São Institute - University of São Federal University of Bahia Paulo Paulo Salvador, Brazil São Carlos, SP – Brazil São Carlos, SP – Brazil [email protected] [email protected] [email protected] Frederico Araújo Durão Department of Computer Science Federal University of Bahia Salvador, Brazil [email protected] ABSTRACT important tools in assisting users to filter what is relevant In this paper, we analyze the application of ensemble al- in this complex information world. There are a number of gorithms to improve the ranking recommendation problem ways to build recommender systems; they are classified as with multiple metadata. We propose three generic ensemble content-based filtering, collaborative filtering or the hybrid strategies that do not require modification of the recom- approach, which combines both filtering strategies [1, 5]. mender algorithm. They combine predictions from a recom- Content-based filtering recommends multimedia content mender trained with distinct metadata into a unified rank to the user based on a profile containing information re- of recommended items. The proposed strategies are Most garding the content, such as genre, keywords, subject, etc. Pleasure, Best of All and Genetic Algorithm Weighting. These metadata are weighted according to past ratings, in The evaluation using the HetRec 2011 MovieLens 2k dataset order to characterize the user's main interests. However, with five different metadata (genres, tags, directors, actors this approach has problems such as over-specialization [1] and countries) shows that our proposed ensemble algorithms and limited performance due to metadata scarcity or quality. achieve a considerable 7% improvement in the Mean Aver- An alternative to this problem is the collaborative filtering, age Precision even with state-of-art collaborative filtering which is based on clusters of similar users or items. One dis- algorithms. advantage of collaborative filtering is the computational ef- fort spent to calculate similarity between users and/or items in a vectorial space composed of user ratings in a user-item Categories and Subject Descriptors matrix. H.3.3 [Information Search and Retrieval]: Information Such limitations have inspired researchers to use matrix Filtering factorization techniques, such as Singular Value Decomposi- tion (SVD), in order to extract latent semantic relationships General Terms between users and items, transforming the vectorial space into a feature space containing topics of interest [20, 11, 17, Design, Algorithms 10]. Nevertheless, other challenges have to be dealt with, such as sparsity, overfitting and data distortion caused by Keywords imputation methods [10]. recommendation; ensemble; metadata; movie; collaborative Considering the limitations and challenges depicted above, filtering hybrid recommenders play an important role because they group together the benefits of content based and collabo- 1. INTRODUCTION rative filtering. It is known that limitations of both ap- proaches, such as the cold start problem, overspecialization Recommender systems have become increasingly popular and limited content analysis, can be reduced when combin- and widely adopted by many sites and services. They are ing both strategies into a unified model [1]. However, most Permission to make digital or hard copies of all or part of this work for recent systems which exploit latent factor models do not con- personal or classroom use is granted without fee provided that copies are not sider the metadata associated to the content, which could made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components provide significant and meaningful information about the of this work owned by others than ACM must be honored. Abstracting with user's interests. Another issue of current metadata aware credit is permitted. To copy otherwise, or republish, to post on servers or to recommenders is that usually they support only one type redistribute to lists, requires prior specific permission and/or a fee. Request of item attribute at a time. To overcome this issue, Bel- permissions from [email protected]. tr~ao et al. [3] analyzed the performance of a recommender WebMedia’14, November 18–21, 2014, João Pessoa, Brazil. using multiple types of metadata, by concatenating the dif- Copyright 2014 ACM 978-1-4503-3230-9/14/11 ...$15.00. http://dx.doi.org/10.1145/2664551.2664569. 231 ferent pieces of information, and although the performance a set of metadata aware algorithms which use the Bayesian improved, the results were still modest. Personalized Ranking (BPR) framework [6] to personalize a Similarly to Beltr~ao et al. [3], this paper proposes a differ- ranking of items using only implicit feedback. These tech- ent approach for handling multiple metadata, using ensem- niques will be considered in our evaluation in the context of ble algorithms. We use three different ensemble strategies movies recommendation. to combine different metadata, but with the advantage that it does not require the algorithm to be modified, or to be 3.1 Notation trained multiple times with the same dataset, and therefore, Following the same notation in [10, 12], we use special it can be used in all current Recommender Systems. indexing letters to distinguish users, items and attributes: This work is structured as follows: in Section 2 we review a user is indicated as u, an item is referred as i; j; k and an related works that use ensemble algorithms; in Section 3 we item's attribute as g. The notation rui is used to refer to briefly describe the models considered in this evaluation; in explicit or implicit feedback from a user u to an item i. In the Section 4 we detail our proposed Ensemble framework and first case, it is an integer provided by the user indicating how strategies; Section 5 presents the evaluation and validation much he liked the content; in the second, it is just a boolean of the approach with HetRec dataset with 855598 ratings, indicating whether the user consumed or visited the content and analysis of the performance of the three proposed strate- or not. The prediction of the system about the preference gies; and finally, in Section 6 we discuss the final remarks, of user u to item i is represented byr ^ui, which is a floating future work and acknowledgments. point value calculated by the recommender algorithm. The set of pairs (u; i) for which rui is known is represented by the set K = f(u; i)jrui is knowng. 2. RELATED WORK Additional sets used in this paper are: N(u) to indicate An ensemble method combines the predictions of different the set of items for which user u provided an implicit feed- algorithms, or the same algorithm with different parameters back, and N¯(u) to indicate the set of items that is unknown to obtain a final prediction. Ensemble algorithms have been to user u. successfully used, for instance, in the Netflix Prize contest consisting of the majority of the top performing solutions. 3.2 BPR-Linear [23, 18]. The BPR-Linear [6] is an algorithm based on the Bayesian Most of the related works in the literature point out that Personalized Ranking (BPR) framework, which uses item ensemble learning has been used in recommender system as attributes in a linear mapping for score estimation. The a way of combining the prediction of multiple algorithms prediction rule is defined as: (heterogeneous ensemble) to create a stronger rank [9], in a technique known as blending. They have been also used n X with a single collaborative filtering algorithm (single-model r^ui = φf (~ai) = wugaig ; (1) or homogeneous ensemble), with methods as Bagging and g=1 Boosting [2]. However, those solutions do not consider the n where φf : R ! R is a function that maps the item at- multiple metadata present in the items, and are often not tributes to the general preferencesr ^ui and ~ai is a boolean practical to implement in a production scenario because of vector of size n where each element aig represents the oc- the computational cost and complexity. In the case of het- currence or not of an attribute, and wug is a weight matrix erogeneous ensemble, it needs to train all models in parallel learned using LearnBPR, which is variation of the stochastic and treat the ensemble as one big model, but unfortunately gradient descent technique [7]. This way, we first compute training 100+ models in parallel and tuning all parameters the relative importance between two items: simultaneously is computationally not feasible [23]. In con- trast, the homogeneous ensemble demands the same model s^uij =r ^ui − r^uj to be trained multiple times, and some methods such as n n Boosting requires that the underlying algorithm be modi- X X = wugaig − wugajg fied to handle the weighted samples. Beltr~ao et al. [3] tried g=1 g=1 (2) n a different approach and combined multiple metadata by X concatenating them, with a modest performance increase. = wug(aig − ajg) : In comparison to the above approaches, our method uses g=1 three different ensemble strategies to combine distinct meta- Finally, the partial derivative with respect to wug is taken: data, but with the advantage that it does not require the al- gorithm to be modified, or to be trained multiple times with @ s^ = (a − a ) ; (3) the same dataset, and therefore, it can be used in all of the @w uij ig jg current Recommender Systems.
Recommended publications
  • The Use of Time Dimension in Recommender Systems for Learning
    The Use of Time Dimension in Recommender Systems for Learning Eduardo José de Borba1, Isabela Gasparini1 and Daniel Lichtnow2 1Graduate Program in Applied Computing (PPGCA), Department of Computer Science (DCC), Santa Catarina State University (UDESC), Paulo Malschitzki 200, Joinville, Brazil 2Polytechnic School, Federal University of Santa Maria (UFSM), Av. Roraima 1000, Santa Maria, Brazil Keywords: Recommender System, Context-aware, Time, Learning. Abstract: When the amount of learning objects is huge, especially in the e-learning context, users could suffer cognitive overload. That way, users cannot find useful items and might feel lost in the environment. Recommender systems are tools that suggest items to users that best match their interests and needs. However, traditional recommender systems are not enough for learning, because this domain needs more personalization for each user profile and context. For this purpose, this work investigates Time-Aware Recommender Systems (Context-aware Recommender Systems that uses time dimension) for learning. Based on a set of categories (defined in previous works) of how time is used in Recommender Systems regardless of their domain, scenarios were defined that help illustrate and explain how each category could be applied in learning domain. As a result, a Recommender System for learning is proposed. It combines Content-Based and Collaborative Filtering approaches in a Hybrid algorithm that considers time in Pre- Filtering and Post-Filtering phases. 1 INTRODUCTION potential to improve the quality of recommendation (Campos et al., 2014). Nowadays, there are distinct educational approaches, This work aims to identify how Time-aware RS e.g. online learning, blended learning, face-to-face (Context-aware RS that uses time context) can be learning, etc.
    [Show full text]
  • A Grouplens Perspective
    From: AAAI Technical Report WS-98-08. Compilation copyright © 1998, AAAI (www.aaai.org). All rights reserved. RecommenderSystems: A GroupLensPerspective Joseph A. Konstan*t , John Riedl *t, AI Borchers,* and Jonathan L. Herlocker* *GroupLensResearch Project *Net Perceptions,Inc. Dept. of ComputerScience and Engineering 11200 West78th Street University of Minnesota Suite 300 Minneapolis, MN55455 Minneapolis, MN55344 http://www.cs.umn.edu/Research/GroupLens/ http://www.netperceptions.com/ ABSTRACT identifying sets of articles by keyworddoes not scale to a In this paper, wereview the history and research findings of situation in which there are thousands of articles that the GroupLensResearch project I and present the four broad contain any imaginable set of keywords. Taken together, research directions that we feel are most critical for these two weaknesses represented an opportunity for a new recommender systems. type of filtering, that would focus on finding which INTRODUCTION:A History of the GroupLensProject available articles matchhuman notions of quality and taste. The GroupLens Research project began at the Computer Such a system would be able to produce a list of articles Supported Cooperative Work (CSCW)Conference in 1992. that each user wouldlike, independentof their content. Oneof the keynote speakers at the conference lectured on a Wedecided to apply our ideas in the domain of Usenet his vision of an emerging information economy,in which news. Usenet screamsfor better information filtering, with most of the effort in the economywould revolve around hundreds of thousands of articles posted daily. Manyof the production, distribution, and consumptionof information, articles in each Usenet newsgroupare on the sametopic, so rather than physical goods and services.
    [Show full text]
  • Evaluation of Recommendation System for Sustainable E-Commerce: Accuracy, Diversity and Customer Satisfaction
    Preprints (www.preprints.org) | NOT PEER-REVIEWED | Posted: 2 January 2020 doi:10.20944/preprints202001.0015.v1 Article Evaluation of Recommendation System for Sustainable E-Commerce: Accuracy, Diversity and Customer Satisfaction Qinglong Li 1, Ilyoung Choi 2 and Jaekyeong Kim 3,* 1 Department of Social Network Science, KyungHee University, 26, Kyungheedae-ro, Dongdaemun-gu, Seoul, 02453, Republic of Korea; [email protected] 2 Graduate School of Business Administration, KyungHee University, 26, Kyungheedae-ro, Dongdaemun- gu, Seoul, 02453, Republic of Korea; [email protected] 3 School of Management, KyungHee University, 26, Kyungheedae-ro, Dongdaemun-gu, Seoul, 02453, Republic of Korea; [email protected] * Correspondence: [email protected]; Tel.: +82-2-961-9355 (F.L.) Abstract: With the development of information technology and the popularization of mobile devices, collecting various types of customer data such as purchase history or behavior patterns became possible. As the customer data being accumulated, there is a growing demand for personalized recommendation services that provide customized services to customers. Currently, global e-commerce companies offer personalized recommendation services to gain a sustainable competitive advantage. However, previous research on recommendation systems has consistently raised the issue that the accuracy of recommendation algorithms does not necessarily lead to the satisfaction of recommended service users. It also claims that customers are highly satisfied when the recommendation system recommends diverse items to them. In this study, we want to identify the factors that determine customer satisfaction when using the recommendation system which provides personalized services. To this end, we developed a recommendation system based on Deep Neural Networks (DNN) and measured the accuracy of recommendation service, the diversity of recommended items and customer satisfaction with the recommendation service.
    [Show full text]
  • Mobile Application Recommender System
    UPTEC IT 10 025 Examensarbete 30 hp December 2010 Mobile Application Recommender System Christoffer Davidsson Abstract Mobile Application Recommender System Christoffer Davidsson Teknisk- naturvetenskaplig fakultet UTH-enheten With the amount of mobile applications available increasing rapidly, users have to put a lot of effort into finding applications of interest. The purpose of this thesis is to Besöksadress: investigate how to aid users in the process of discovering new mobile applications by Ångströmlaboratoriet Lägerhyddsvägen 1 providing them with recommendations. A prototype system is then built as a Hus 4, Plan 0 proof-of-concept. Postadress: The work of the thesis is divided into three phases where the aim of the first phase is Box 536 751 21 Uppsala to study related work and related systems to identify promising concepts and features. During the second phase, a prototype system is designed and implemented. Telefon: The outcome and result of the first two phases is then evaluated and analyzed in the 018 – 471 30 03 third and final phase. Telefax: 018 – 471 30 00 The prototype system integrates and extends an existing recommender engine previously used to recommend media items. As a part of the system, an Android Hemsida: application is developed, which observes user actions and presents recommended http://www.teknat.uu.se/student applications to the user. In parallel to the development, the system was tested by a small group of users recruited among colleagues at Ericsson. The data generated during this test period is analyzed to show the usefulness of observed user actions over explicit ratings and the dependency on context for application usage.
    [Show full text]
  • A Systematic Review and Taxonomy of Explanations in Decision Support and Recommender Systems
    Noname manuscript No. (will be inserted by the editor) A Systematic Review and Taxonomy of Explanations in Decision Support and Recommender Systems Ingrid Nunes · Dietmar Jannach Received: date / Accepted: date Abstract With the recent advances in the field of artificial intelligence, an increasing number of decision-making tasks are delegated to software systems. A key requirement for the success and adoption of such systems is that users must trust system choices or even fully automated decisions. To achieve this, explanation facilities have been widely investigated as a means of establishing trust in these systems since the early years of expert systems. With today's increasingly sophisticated machine learning algorithms, new challenges in the context of explanations, accountability, and trust towards such systems con- stantly arise. In this work, we systematically review the literature on expla- nations in advice-giving systems. This is a family of systems that includes recommender systems, which is one of the most successful classes of advice- giving software in practice. We investigate the purposes of explanations as well as how they are generated, presented to users, and evaluated. As a result, we derive a novel comprehensive taxonomy of aspects to be considered when de- signing explanation facilities for current and future decision support systems. The taxonomy includes a variety of different facets, such as explanation objec- tive, responsiveness, content and presentation. Moreover, we identified several challenges that remain unaddressed so far, for example related to fine-grained issues associated with the presentation of explanations and how explanation facilities are evaluated. Keywords Explanation · Decision Support System · Recommender System · Expert System · Knowledge-based System · Systematic Review · Machine Learning · Trust · Artificial Intelligence I.
    [Show full text]
  • Collaborative Filtering: a Machine Learning Perspective by Benjamin Marlin a Thesis Submitted in Conformity with the Requirement
    Collaborative Filtering: A Machine Learning Perspective by Benjamin Marlin A thesis submitted in conformity with the requirements for the degree of Master of Science Graduate Department of Computer Science University of Toronto Copyright c 2004 by Benjamin Marlin Abstract Collaborative Filtering: A Machine Learning Perspective Benjamin Marlin Master of Science Graduate Department of Computer Science University of Toronto 2004 Collaborative filtering was initially proposed as a framework for filtering information based on the preferences of users, and has since been refined in many different ways. This thesis is a comprehensive study of rating-based, pure, non-sequential collaborative filtering. We analyze existing methods for the task of rating prediction from a machine learning perspective. We show that many existing methods proposed for this task are simple applications or modifications of one or more standard machine learning methods for classification, regression, clustering, dimensionality reduction, and density estima- tion. We introduce new prediction methods in all of these classes. We introduce a new experimental procedure for testing stronger forms of generalization than has been used previously. We implement a total of nine prediction methods, and conduct large scale prediction accuracy experiments. We show interesting new results on the relative performance of these methods. ii Acknowledgements I would like to begin by thanking my supervisor Richard Zemel for introducing me to the field of collaborative filtering, for numerous helpful discussions about a multitude of models and methods, and for many constructive comments about this thesis itself. I would like to thank my second reader Sam Roweis for his thorough review of this thesis, as well as for many interesting discussions of this and other research.
    [Show full text]
  • User Data Analytics and Recommender System for Discovery Engine
    User Data Analytics and Recommender System for Discovery Engine Yu Wang Master of Science Thesis Stockholm, Sweden 2013 TRITA-ICT-EX-2013: 88 User Data Analytics and Recommender System for Discovery Engine Yu Wang [email protected] Whaam AB, Stockholm, Sweden Royal Institute of Technology, Stockholm, Sweden June 11, 2013 Supervisor: Yi Fu, Whaam AB Examiner: Prof. Johan Montelius, KTH Abstract On social bookmarking website, besides saving, organizing and sharing web pages, users can also discovery new web pages by browsing other’s bookmarks. However, as more and more contents are added, it is hard for users to find interesting or related web pages or other users who share the same interests. In order to make bookmarks discoverable and build a discovery engine, sophisticated user data analytic methods and recommender system are needed. This thesis addresses the topic by designing and implementing a prototype of a recommender system for recommending users, links and linklists. Users and linklists recommendation is calculated by content- based method, which analyzes the tags cosine similarity. Links recommendation is calculated by linklist based collaborative filtering method. The recommender system contains offline and online subsystem. Offline subsystem calculates data statistics and provides recommendation candidates while online subsystem filters the candidates and returns to users. The experiments show that in social bookmark service like Whaam, tag based cosine similarity method improves the mean average precision by 45% compared to traditional collaborative method for user and linklist recommendation. For link recommendation, the linklist based collaborative filtering method increase the mean average precision by 39% compared to the user based collaborative filtering method.
    [Show full text]
  • A Recommender System for Groups of Users
    PolyLens: A Recommender System for Groups of Users Mark O’Connor, Dan Cosley, Joseph A. Konstan, and John Riedl Department of Computer Science and Engineering, University of Minnesota, Minneapolis, MN 55455 USA {oconnor;cosley;konstan;riedl}@cs.umn.edu Abstract. We present PolyLens, a new collaborative filtering recommender system designed to recommend items for groups of users, rather than for individuals. A group recommender is more appropriate and useful for domains in which several people participate in a single activity, as is often the case with movies and restaurants. We present an analysis of the primary design issues for group recommenders, including questions about the nature of groups, the rights of group members, social value functions for groups, and interfaces for displaying group recommendations. We then report on our PolyLens prototype and the lessons we learned from usage logs and surveys from a nine-month trial that included 819 users. We found that users not only valued group recommendations, but were willing to yield some privacy to get the benefits of group recommendations. Users valued an extension to the group recommender system that enabled them to invite non-members to participate, via email. Introduction Recommender systems (Resnick & Varian, 1997) help users faced with an overwhelming selection of items by identifying particular items that are likely to match each user’s tastes or preferences (Schafer et al., 1999). The most sophisticated systems learn each user’s tastes and provide personalized recommendations. Though several machine learning and personalization technologies can attempt to learn user preferences, automated collaborative filtering (Resnick et al., 1994; Shardanand & Maes, 1995) has become the preferred real-time technology for personal recommendations, in part because it leverages the experiences of an entire community of users to provide high quality recommendations without detailed models of either content or user tastes.
    [Show full text]
  • Chapter: Music Recommender Systems
    Chapter 13 Music Recommender Systems Markus Schedl, Peter Knees, Brian McFee, Dmitry Bogdanov, and Marius Kaminskas 13.1 Introduction Boosted by the emergence of online music shops and music streaming services, digital music distribution has led to an ubiquitous availability of music. Music listeners, suddenly faced with an unprecedented scale of readily available content, can easily become overwhelmed. Music recommender systems, the topic of this chapter, provide guidance to users navigating large collections. Music items that can be recommended include artists, albums, songs, genres, and radio stations. In this chapter, we illustrate the unique characteristics of the music recommen- dation problem, as compared to other content domains, such as books or movies. To understand the differences, let us first consider the amount of time required for ausertoconsumeasinglemediaitem.Thereisobviouslyalargediscrepancyin consumption time between books (days or weeks), movies (one to a few hours), and asong(typicallyafewminutes).Consequently,thetimeittakesforausertoform opinions for music can be much shorter than in other domains, which contributes to the ephemeral, even disposable, nature of music. Similarly, in music, a single M. Schedl (!)•P.Knees Department of Computational Perception, Johannes Kepler University Linz, Linz, Austria e-mail: [email protected]; [email protected] B. McFee Center for Data Science, New York University, New York, NY, USA e-mail: [email protected] D. Bogdanov Music Technology Group, Universitat Pompeu Fabra, Barcelona, Spain e-mail: [email protected] M. Kaminskas Insight Centre for Data Analytics, University College Cork, Cork, Ireland e-mail: [email protected] ©SpringerScience+BusinessMediaNewYork2015 453 F. Ricci et al. (eds.), Recommender Systems Handbook, DOI 10.1007/978-1-4899-7637-6_13 454 M.
    [Show full text]
  • A Multitask Ranking System
    Recommending What Video to Watch Next: A Multitask Ranking System Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, Ed Chi Google, Inc. {zhezhao,lichan,liwei,jilinc,aniruddhnath,shawnandrews,aditeek,nlogn,xinyang,edchi}@google.com ABSTRACT promising items. We present experiments and lessons learned from In this paper, we introduce a large scale multi-objective ranking building such a ranking system on a large-scale industrial video system for recommending what video to watch next on an indus- publishing and sharing platform. trial video sharing platform. The system faces many real-world Designing and developing a real-world large-scale video recom- challenges, including the presence of multiple competing ranking mendation system is full of challenges, including: objectives, as well as implicit selection biases in user feedback. To • There are often diferent and sometimes conficting objec- tackle these challenges, we explored a variety of soft-parameter tives which we want to optimize for. For example, we may sharing techniques such as Multi-gate Mixture-of-Experts so as to want to recommend videos that users rate highly and share efciently optimize for multiple ranking objectives. Additionally, with their friends, in addition to watching. we mitigated the selection biases by adopting a Wide & Deep frame- • There is often implicit bias in the system. For example, a user work. We demonstrated that our proposed techniques can lead to might have clicked and watched a video simply because it substantial improvements on recommendation quality on one of was being ranked high, not because it was the one that the the world’s largest video sharing platforms.
    [Show full text]
  • When Diversity Met Accuracy: a Story of Recommender Systems †
    proceedings Extended Abstract When Diversity Met Accuracy: A Story of Recommender Systems † Alfonso Landin * , Eva Suárez-García and Daniel Valcarce Department of Computer Science, University of A Coruña, 15071 A Coruña, Spain; [email protected] (E.S.-G.); [email protected] (D.V.) * Correspondence: [email protected]; Tel.: +34-881-01-1276 † Presented at the XoveTIC Congress, A Coruña, Spain, 27–28 September 2018. Published: 14 September 2018 Abstract: Diversity and accuracy are frequently considered as two irreconcilable goals in the field of Recommender Systems. In this paper, we study different approaches to recommendation, based on collaborative filtering, which intend to improve both sides of this trade-off. We performed a battery of experiments measuring precision, diversity and novelty on different algorithms. We show that some of these approaches are able to improve the results in all the metrics with respect to classical collaborative filtering algorithms, proving to be both more accurate and more diverse. Moreover, we show how some of these techniques can be tuned easily to favour one side of this trade-off over the other, based on user desires or business objectives, by simply adjusting some of their parameters. Keywords: recommender systems; collaborative filtering; diversity; novelty 1. Introduction Over the years the user experience with different services has shifted from a proactive approach, where the user actively look for content, to one where the user is more passive and content is suggested to her by the service. This has been possible due to the advance in the field of recommender systems (RS), making it possible to make better suggestions to the users, personalized to their preferences.
    [Show full text]
  • Indian Regional Movie Dataset for Recommender Systems
    Indian Regional Movie Dataset for Recommender Systems Prerna Agarwal∗ Richa Verma∗ Angshul Majumdar IIIT Delhi IIIT Delhi IIIT Delhi [email protected] [email protected] [email protected] ABSTRACT few years with a lot of diversity in languages and viewers. As per Indian regional movie dataset is the rst database of regional Indian the UNESCO cinema statistics [9], India produces around 1,724 movies, users and their ratings. It consists of movies belonging to movies every year with as many as 1,500 movies in Indian regional 18 dierent Indian regional languages and metadata of users with languages. India’s importance in the global lm industry is largely varying demographics. rough this dataset, the diversity of Indian because India is home to Bollywood in Mumbai. ere’s a huge base regional cinema and its huge viewership is captured. We analyze of audience in India with a population of 1.3 billion which is evident the dataset that contains roughly 10K ratings of 919 users and by the fact that there are more than two thousand multiplexes in 2,851 movies using some supervised and unsupervised collaborative India where over 2.2 billion movie tickets were sold in 2016. e box ltering techniques like Probabilistic Matrix Factorization, Matrix oce revenue in India keeps on rising every year. erefore, there Completion, Blind Compressed Sensing etc. e dataset consists of is a huge need for a dataset like Movielens in Indian context that can metadata information of users like age, occupation, home state and be used for testing and bench-marking recommendation systems known languages.
    [Show full text]