Enhance Nmf-Based Recommendation Systems with Auxiliary Information Imputation

Total Page:16

File Type:pdf, Size:1020Kb

Enhance Nmf-Based Recommendation Systems with Auxiliary Information Imputation University of Kentucky UKnowledge Theses and Dissertations--Computer Science Computer Science 2019 ENHANCE NMF-BASED RECOMMENDATION SYSTEMS WITH AUXILIARY INFORMATION IMPUTATION Fatemah Alghamedy University of Kentucky, [email protected] Digital Object Identifier: https://doi.org/10.13023/etd.2019.144 Right click to open a feedback form in a new tab to let us know how this document benefits ou.y Recommended Citation Alghamedy, Fatemah, "ENHANCE NMF-BASED RECOMMENDATION SYSTEMS WITH AUXILIARY INFORMATION IMPUTATION" (2019). Theses and Dissertations--Computer Science. 79. https://uknowledge.uky.edu/cs_etds/79 This Doctoral Dissertation is brought to you for free and open access by the Computer Science at UKnowledge. It has been accepted for inclusion in Theses and Dissertations--Computer Science by an authorized administrator of UKnowledge. For more information, please contact [email protected]. STUDENT AGREEMENT: I represent that my thesis or dissertation and abstract are my original work. Proper attribution has been given to all outside sources. I understand that I am solely responsible for obtaining any needed copyright permissions. I have obtained needed written permission statement(s) from the owner(s) of each third-party copyrighted matter to be included in my work, allowing electronic distribution (if such use is not permitted by the fair use doctrine) which will be submitted to UKnowledge as Additional File. I hereby grant to The University of Kentucky and its agents the irrevocable, non-exclusive, and royalty-free license to archive and make accessible my work in whole or in part in all forms of media, now or hereafter known. I agree that the document mentioned above may be made available immediately for worldwide access unless an embargo applies. I retain all other ownership rights to the copyright of my work. I also retain the right to use in future works (such as articles or books) all or part of my work. I understand that I am free to register the copyright to my work. REVIEW, APPROVAL AND ACCEPTANCE The document mentioned above has been reviewed and accepted by the student’s advisor, on behalf of the advisory committee, and by the Director of Graduate Studies (DGS), on behalf of the program; we verify that this is the final, approved version of the student’s thesis including all changes required by the advisory committee. The undersigned agree to abide by the statements above. Fatemah Alghamedy, Student Dr. Jun Zhang, Major Professor Dr. Miroslaw Truszczynski, Director of Graduate Studies ENHANCE NMF-BASED RECOMMENDATION SYSTEMS WITH AUXILIARY INFORMATION IMPUTATION |||||||||||||| DISSERTATION |||||||||||||| A dissertation submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy in the College of Engineering at the University of Kentucky By Fatemah Alghamedy Lexington, Kentucky Director: Jun Zhang, Ph.D., Professor of Computer Science Lexington, Kentucky 2019 Copyright c Fatemah Alghamedy 2019 ABSTRACT OF DISSERTATION ENHANCE NMF-BASED RECOMMENDATION SYSTEMS WITH AUXILIARY INFORMATION IMPUTATION This dissertation studies the factors that negatively impact the accuracy of the collaborative filtering recommendation systems based on nonnegative matrix factorization (NMF). The keystone in the recommendation system is the rating that expresses the user's opinion about an item. One of the most significant issues in the recommendation systems is the lack of ratings. This issue is called "cold- start" issue, which appears clearly with New-Users who did not rate any item and New-Items, which did not receive any rating. The traditional recommendation systems assume that users are independent and identically distributed and ignore the connections among users whereas the recommendation actually is a social activity. This dissertation aims to enhance NMF-based recommendation systems by utilizing the imputation method and lim- iting the errors that are introduced in the system. External information such as trust network and item categories are incorporated into NMF-based recommen- dation systems through the imputation. The proposed approaches impute various subsets of the missing ratings. The subsets are defined based on the total number of the ratings of the user or item before the imputation, such as impute the missing ratings of New-Users, New- Items, or cold-start users or items that suffer from the lack of the ratings. In addition, several factors are analyzed that affect the prediction accuracy when the imputation method is utilized with NMF-based recommendation systems. These factors include the total number of the ratings of the user or item before the im- putation, the total number of imputed ratings for each user and item, the average of imputed rating values, and the value of imputed rating values. In addition, several strategies are applied to select the subset of missing ratings for the impu- tation that lead to increasing the prediction accuracy and limiting the imputation error. Moreover, a comparison is conducted with some popular methods that are in common with the proposed method in utilizing the imputation to handle the lack of ratings, but they differ in the source of the imputed ratings. Experiments on different large-size datasets are conducted to examine the pro- posed approaches and analyze the effects of the imputation on accuracy. Users and items are divided into three groups based on the total number of the ratings before the imputation is applied and their recommendation accuracy is calculated. The results show that the imputation enhances the recommendation system by capacitating the system to recommend items to New-Users, introduce New-Items to users, and increase the accuracy of the cold-start users and items. However, the analyzed factors play important roles in the recommendation accuracy and limit the error that is introduced from the imputation. KEYWORDS: recommendation system, collaborative filtering, NMF, trust ma- trix, imputation. Fatemah Alghamedy ||||||||||||{ April 16, 2019 ||||||||||||{ ENHANCE NMF-BASED RECOMMENDATION SYSTEMS WITH AUXILIARY INFORMATION IMPUTATION By Fatemah Alghamedy Jun Zhang, Ph.D. ||||||||||||{ Director of Dissertation Miroslaw Truszczynski, Ph.D. ||||||||||||{ Director of Graduate Studies April 16, 2019 ||||||||||||{ Date To my father, my mother, my husband, my daughter, my son, my sisters, and my brother. ACKNOWLEDGMENTS I would like to express my gratitude to the people who helped and encouraged me during my Ph.D. journey. First of all, I would like to express the deepest appreciation to my academic advisor, Prof. Jun Zhang, for his continuous support, patience, motivation, en- thusiasm, and immense knowledge. His guidance helped me throughout my time of researching and writing this dissertation. Next, I would like to thank the committee faculty members Dr. Jerzy W. Jaromczyk (Department of Computer Science), Dr. Jinze Liu (Department of Computer Science), Dr. Sen-ching Samson Cheung (Department of Electrical and Computer Engineering), and the outside examiner Prof. Qiang Ye (Department of Mathematics) for their time and the helpful comments on my dissertation. Also, I would like to express my special thanks to the Computer Science De- partment and to UK Research Computing (UKRC) for their research resources. In addition, I would like to thank the Imam Abdulrahman bin Faisal University and Saudi Arabian Cultural Mission (SACM) for their full financial support for the entirety of my Ph.D. years. Last but not least, I would like to thank my family for their unlimited support, especially my father and my mother, whom I wouldn't be able to reach this success without their care and support ever since I was a child. My father taught me the meaning of determination and ambition and was always there when I needed any advice or opinion, especially concerning the academic field. My mother surrounded me by her love and care and never forgot to support me and take care of me even though there were thousands of miles distance between us. A special thanks to my adorable children, Aseel and Suhail. Thank you for bearing with me on this journey, as I'm mostly busy; and above all of that, thank you for being a source of strength for me. Having you around me gave me the will to continue and to keep moving forward. iii TABLE OF CONTENTS Acknowledgments iii List of Tables vii List of Figures ix 1 Introduction 1 1.1 Related Works . .5 1.2 Dissertation Organization . 10 1.2.1 Technical Contributions . 10 1.2.2 Notational Conventions . 11 1.2.3 Data Description . 11 1.2.4 Evaluation Strategy . 14 2 Imputation with Item Auxiliary Information 16 2.1 Problem Description . 16 2.2 Proposed Method . 19 2.2.1 Objective Function . 21 2.2.2 Update Formula . 21 2.2.3 Convergence Analysis . 22 2.2.4 Detailed Algorithm . 23 2.2.5 Complexity . 25 2.3 Experimental Study . 28 2.3.1 The Influence of Imputed Rating Value . 30 2.3.2 Parameter Study . 34 2.4 Summary . 36 3 Imputation with Trust Network Information for New Users 38 3.1 Problem Description . 38 3.2 Proposed Method . 41 3.2.1 Objective Function . 41 3.2.2 Update Formula . 42 3.2.3 Convergence Analysis . 42 3.2.4 Detailed Algorithm . 42 3.2.5 Complexity . 44 3.3 Experimental Study . 45 3.3.1 Imputation Vs. WNMTF . 48 3.4 Summary . 50 4 Influential Factors of Imputation with Trust Network Information for Cold-Start Users 51 4.1 Problem Description . 51 4.2 Proposed Method . 52 4.2.1 Objective Function . 54 4.2.2 Update Formula . 55 iv 4.2.3 Convergence Analysis . 55 4.2.4 Detailed Algorithm . 55 4.2.5 Complexity . 57 4.3 Experimental Study . 57 4.3.1 Influence of the Rating Value Average . 62 4.3.2 Parameter Settings . 63 4.3.3 Summary of Results . 66 4.4 Summary . 67 5 Selective Imputation Strategies Based on Fused Factored Matri- ces 68 5.1 Problem Description . 68 5.2 Proposed Method .
Recommended publications
  • Imputets: Time Series Missing Value Imputation in R by Steffen Moritz and Thomas Bartz-Beielstein
    CONTRIBUTED RESEARCH ARTICLE 1 imputeTS: Time Series Missing Value Imputation in R by Steffen Moritz and Thomas Bartz-Beielstein Abstract The imputeTS package specializes on univariate time series imputation. It offers multiple state-of-the-art imputation algorithm implementations along with plotting functions for time series missing data statistics. While imputation in general is a well-known problem and widely covered by R packages, finding packages able to fill missing values in univariate time series is more complicated. The reason for this lies in the fact, that most imputation algorithms rely on inter-attribute correlations, while univariate time series imputation instead needs to employ time dependencies. This paper provides an introduction to the imputeTS package and its provided algorithms and tools. Furthermore, it gives a short overview about univariate time series imputation in R. Introduction In almost every domain from industry (Billinton et al., 1996) to biology (Bar-Joseph et al., 2003), finance (Taylor, 2007) up to social science (Gottman, 1981) different time series data are measured. While the recorded datasets itself may be different, one common problem are missing values. Many analysis methods require missing values to be replaced with reasonable values up-front. In statistics this process of replacing missing values is called imputation. Time series imputation thereby is a special sub-field in the imputation research area. Most popular techniques like Multiple Imputation (Rubin, 1987), Expectation-Maximization (Dempster et al., 1977), Nearest Neighbor (Vacek and Ashikaga, 1980) and Hot Deck (Ford, 1983) rely on inter-attribute correlations to estimate values for the missing data. Since univariate time series do not possess more than one attribute, these algorithms cannot be applied directly.
    [Show full text]
  • Uncertainty-Aware Variational-Recurrent Imputation Network for Clinical Time Series Ahmad Wisnu Mulyadi, Eunji Jun, and Heung-Il Suk, Member, IEEE
    UNDER REVIEW 1 Uncertainty-Aware Variational-Recurrent Imputation Network for Clinical Time Series Ahmad Wisnu Mulyadi, Eunji Jun, and Heung-Il Suk, Member, IEEE Abstract—Electronic health records (EHR) consist of longitudi- data yields poor performance at high missing rates and also nal clinical observations portrayed with sparsity, irregularity, and requires modeling separately for the distinct dataset. In fact, high-dimensionality, which become major obstacles in drawing those missing values reflect the decisions made by health-care reliable downstream clinical outcomes. Although there exist great numbers of imputation methods to tackle these issues, most of providers [5]. Therefore, the missing values and their patterns them ignore correlated features, temporal dynamics and entirely contain information regarding a patient’s health status [5] and set aside the uncertainty. Since the missing value estimates involve correlate with the outcomes or target labels [6]. Thus, we the risk of being inaccurate, it is appropriate for the method to resort to the imputation approach by exploiting those missing handle the less certain information differently than the reliable patterns to improve the prediction of the clinical outcomes as data. In that regard, we can use the uncertainties in estimating the missing values as the fidelity score to be further utilized to the downstream task. alleviate the risk of biased missing value estimates. In this work, There exist numerous strategies for imputing missing values we propose a novel variational-recurrent imputation network, in the literature. Generally, imputation methods can be clas- which unifies an imputation and a prediction network by taking sified into a deterministic or stochastic approach, depending into account the correlated features, temporal dynamics, as well on the use of randomness [7].
    [Show full text]
  • Potential Sources of Missing Data in a Meta-Analysis
    SMG advanced workshop, Cardiff, March 4-5, 2010 PotentialPotential sourcessources ofof missingmissing datadata inin aa meta-analysismeta-analysis Missing data • Studies not found • Outcome not reported Julian Higgins • Outcome partially reported (e.g. missing SD or SE) MRC Biostatistics Unit Cambridge, UK • Extra information required for meta-analysis (e.g. imputed correlation coefficient) with thanks to Ian White, Fred Wolf, Angela Wood, Alex Sutton • Missing participants • No information on study characteristic for heterogeneity analysis ConceptsConcepts inin missingmissing datadata ConceptsConcepts inin missingmissing datadata 1. ‘Missing completely at random’ 2. ‘Missing at random’ • As if missing observation was randomly picked • Missingness depends on things you know about, and • Missingness unrelated to the true (missing) value not on the missing data themselves • e.g. – (genuinely) accidental deletion of a file •e.g. – observations measured on a random sample – older people more likely to drop out (irrespective of – statistics not reported due to ignorance of their the effect of treatment) importance(?) – recent cluster trials more likely to report intraclass correlation • OK (unbiased) to analyse just the data available, but – experimental treatment more likely to cause drop- sample size is reduced - which may not be ideal out (independent of therapeutic effect) • Not very common! • Usually surmountable StrategiesStrategies forfor dealingdealing withwith missingmissing datadata ConceptsConcepts inin missingmissing datadata •
    [Show full text]
  • Multiple Imputation: a Review of Practical and Theoretical Findings Jared S
    Submitted to Statistical Science Multiple Imputation: A Review of Practical and Theoretical Findings Jared S. Murray∗ University of Texas at Austin Abstract. Multiple imputation is a straightforward method for handling missing data in a principled fashion. This paper presents an overview of multiple imputation, including important theoretical results and their practical implications for generating and using multiple imputations. A review of strategies for generating imputations follows, including recent developments in flexible joint modeling and sequential regres- sion/chained equations/fully conditional specification approaches. Fi- nally, we compare and contrast different methods for generating impu- tations on a range of criteria before identifying promising avenues for future research. Key words and phrases: missing data, proper imputation, congeniality, chained equations, fully conditional specification, sequential regression multivariate imputation. 1. INTRODUCTION Multiple imputation (MI) (Rubin, 1987) is a simple but powerful method for dealing with missing data. MI as originally conceived proceeds in two stages: A data disseminator creates a small number of completed datasets by filling in the missing values with samples from an imputation model. Analysts compute their estimates in each completed dataset and combine them using simple rules to get pooled estimates and standard errors that incorporate the additional variability due to the missing data. MI was originally developed for settings in which statistical agencies or other data disseminators provide multiply imputed databases to distinct end-users. arXiv:1801.04058v1 [stat.ME] 12 Jan 2018 There are a number of benefits to MI in this setting: The disseminator can support approximately valid inference for a wide range of potential analyses with a small set of imputations, and the burden of dealing with the missing data is on the imputer rather than the analyst.
    [Show full text]
  • Within-Industry Productivity Dispersion and Imputation for Missing Data
    Imputation in U.S. Manufacturing Data and Implications for Within-Industry Productivity Dispersion ∗ T. Kirk Whitey, Jerome P. Reiter**, and Amil Petrin*** y Center for Economic Studies, U.S. Census Bureau ** Duke University *** University of Minnesota,Twin Cities and NBER Abstract In the U.S. Census Bureau’s 2002 and 2007 manufacturing data, respectively, 79% and 73% of observa­ tions have imputed data for at least one variable used to compute total factor productivity. The Bureau imputes for missing values using methods known to result in underestimation of variability and potential bias in multivariate inferences. We present an alternative strategy for handling the missing data based on multiple imputation via sequences of classification and regression trees. We use our imputations and the Bureau’s imputations to estimate within-industry productivity dispersions for every manufacturing industry. The results suggest that there is even more within-industry productivity dispersion in U.S. man­ ufacturing than previous research has indicated. We also estimate relationships between plant exit, pro­ ductivity, prices, and demand shocks. For these estimands, we find the results are substantively robust to using an alternative imputation strategy. ∗Some of the research in this paper was conducted while the first author was a Census Bureau employee. Any opinions and conclusions expressed herein are those of the authors and do not necessarily represent the views of the U.S. Census Bureau. All results have been reviewed to ensure that no confidential information is disclosed. Corresponding author: T. Kirk White. Email: [email protected]. 1 1 Introduction Nearly all economic surveys suffer from item nonresponse, i.e., respondents answer some questions but not others.
    [Show full text]
  • Dealing with Missing Standard Deviation and Mean Values in Meta-Analysis of Continuous Outcomes: a Systematic Review Christopher J
    Weir et al. BMC Medical Research Methodology (2018) 18:25 https://doi.org/10.1186/s12874-018-0483-0 RESEARCHARTICLE Open Access Dealing with missing standard deviation and mean values in meta-analysis of continuous outcomes: a systematic review Christopher J. Weir1*, Isabella Butcher1, Valentina Assi1, Stephanie C. Lewis1, Gordon D. Murray1, Peter Langhorne2 and Marian C. Brady3 Abstract Background: Rigorous, informative meta-analyses rely on availability of appropriate summary statistics or individual participant data. For continuous outcomes, especially those with naturally skewed distributions, summary information on the mean or variability often goes unreported. While full reporting of original trial data is the ideal, we sought to identify methods for handling unreported mean or variability summary statistics in meta-analysis. Methods: We undertook two systematic literature reviews to identify methodological approaches used to deal with missing mean or variability summary statistics. Five electronic databases were searched, in addition to the Cochrane Colloquium abstract books and the Cochrane Statistics Methods Group mailing list archive. We also conducted cited reference searching and emailed topic experts to identify recent methodological developments. Details recorded included the description of the method, the information required to implement the method, any underlying assumptions and whether the method could be readily applied in standard statistical software. We provided a summary description of the methods identified, illustrating selected methods in example meta-analysis scenarios. Results: For missing standard deviations (SDs), following screening of 503 articles, fifteen methods were identified in addition to those reported in a previous review. These included Bayesian hierarchical modelling at the meta-analysis level; summary statistic level imputation based on observed SD values from other trials in the meta-analysis; a practical approximation based on the range; and algebraic estimation of the SD based on other summary statistics.
    [Show full text]
  • Missing Value Imputation Using Feature-Specific Generative
    IFGAN: Missing Value Imputation using Feature-specific Generative Adversarial Networks Wei Qiu* Yangsibo Huang* Quanzheng Li Yuanpei College College of Computer Science and Technology Department of Radiology Peking University Zhejiang University Massachusetts General Hospital Beijing, China Hangzhou, China Boston, MA [email protected] [email protected] [email protected] Abstract—Missing value imputation is a challenging and well- established state-of-the-art performance in many applications. researched topic in data mining. In this paper, we propose When it comes to the task of imputation, deep architectures IFGAN, a missing value imputation algorithm based on Feature- have also shown greater promise than classical ones in recent specific Generative Adversarial Networks (GAN). Our idea is intuitive yet effective: a feature-specific generator is trained studies [7]–[9], due to their capability to automatically learn to impute missing values, while a discriminator is expected to latent representations and inter-variable associations. Nonethe- distinguish the imputed values from observed ones. The proposed less, these work all formulated data imputation as a mapping architecture is capable of handling different data types, data from a missing matrix to a complete one, which ignores the distributions, missing mechanisms, and missing rates. It also difference in data type and statistical importance between improves post-imputation analysis by preserving inter-feature correlations. We empirically show on several real-life datasets features and makes the model unable to learn feature-specific that IFGAN outperforms current state-of-the-art algorithm un- representations. der various missing conditions. To overcome the limitations of classical models and to Index Terms—Data preprocessing, Missing value imputation, incorporate feature-specific representations into deep learning Deep adversarial network frameworks, we propose IFGAN, a feature-specific generative adversarial network for missing data imputation.
    [Show full text]
  • Should a Normal Imputation Model Be Modified to Impute Skewed Variables?
    Sociological Methods and Research, 2013, 42(1), 105-138 Should a Normal Imputation Model Be Modified to Impute Skewed Variables? Paul T. von Hippel Abstract (169 words) Researchers often impute continuous variables under an assumption of normality, yet many incomplete variables are skewed. We find that imputing skewed continuous variables under a normal model can lead to bias; the bias is usually mild for popular estimands such as means, standard deviations, and linear regression coefficients, but the bias can be severe for more shape- dependent estimands such as percentiles or the coefficient of skewness. We test several methods for adapting a normal imputation model to accommodate skewness, including methods that transform, truncate, or censor (round) normally imputed values, as well as methods that impute values from a quadratic or truncated regression. None of these modifications reliably reduces the biases of the normal model, and some modifications can make the biases much worse. We conclude that, if one has to impute a skewed variable under a normal model, it is usually safest to do so without modifications—unless you are more interested in estimating percentiles and shape that in estimated means, variance, and regressions. In the conclusion, we briefly discuss promising developments in the area of continuous imputation models that do not assume normality. Key words: missing data; missing values; incomplete data; regression; transformation; normalization; multiple imputation; imputation Paul T. von Hippel is Assistant Professor, LBJ School of Public Affairs, University of Texas, 2315 Red River, Box Y, Austin, TX 78712, [email protected]. I thank Ray Koopman for help with the calculations in Section 2.2.
    [Show full text]
  • Supervised Nonnegative Matrix Factorization to Predict ICU
    Supervised Nonnegative Matrix Factorization to Predict ICU Mortality Risk Guoqing Chao Chengsheng Mao Fei Wang Yuan Zhao Feinberg School of Medicine Feinberg School of Medicine Weill Cornell Medicine Feinberg School of Medicine Northwestern University Northwestern University Cornell University Northwestern University Chicago, U.S Chicago, U.S New York, U.S Chicago, U.S [email protected] [email protected] [email protected] [email protected] Yuan Luo Feinberg School of Medicine Northwestern University Chicago, U.S [email protected] Abstract—ICU mortality risk prediction is a tough yet im- In this paper, we describe a supervised nonnegative ma- portant task. On one hand, due to the complex temporal data trix factorization (NMF) algorithm that performs predictive collected, it is difficult to identify the effective features and modeling by the exploration of atomic and the higher-order interpret them easily; on the other hand, good prediction can help clinicians take timely actions to prevent the mortality. These features jointly. We automate the mining of higher-order fea- correspond to the interpretability and accuracy problems. Most tures by converting data into graph representation. Moreover, existing methods lack of the interpretability, but recently Sub- supervised latent group identification reduces dimensionality graph Augmented Nonnegative Matrix Factorization (SANMF) of different feature types for patients, and simultaneously has been successfully applied to time series data to provide a group temporal trends to form effective features for patient path to interpret the features well. Therefore, we adopted this approach as the backbone to analyze the patient data. One outcome prediction. Applications on patient physiological time limitation of the original SANMF method is its poor prediction series [1] show significant performance improvements from ability due to its unsupervised nature.
    [Show full text]
  • Optimal Recovery of Missing Values for Non-Negative Matrix Factorization Rebecca Chen and Lav R
    bioRxiv preprint doi: https://doi.org/10.1101/647560; this version posted May 24, 2019. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under aCC-BY-NC-ND 4.0 International license. 1 Optimal Recovery of Missing Values for Non-negative Matrix Factorization Rebecca Chen and Lav R. Varshney Abstract We extend the approximation-theoretic technique of optimal recovery to the setting of imputing missing values in clustered data, specifically for non-negative matrix factorization (NMF), and develop an implementable algorithm. Under certain geometric conditions, we prove tight upper bounds on NMF relative error, which is the first bound of this type for missing values. We also give probabilistic bounds for the same geometric assumptions. Experiments on image data and biological data show that this theoretically-grounded technique performs as well as or better than other imputation techniques that account for local structure. Index Terms imputation, missing values, non-negative matrix factorization, optimal recovery I. INTRODUCTION ATRIX factorization is commonly used for clustering and dimensionality reduction in computational biology, imaging, M and other fields. Non-negative matrix factorization (NMF) is particularly favored by biologists because non-negativity constraints preclude negative values that are difficult to interpret in biological processes [2], [3]. NMF of gene expression matrices can discover cell groups and lower-dimensional manifolds in gene counts for different cell types. Due to physical and biological limitations of DNA- and RNA-sequencing techniques, however, gene-expression matrices are usually incomplete and matrix imputation is often necessary before further analysis [3].
    [Show full text]
  • Convex Nonnegative Matrix Factorization with Missing Data Ronan Hamon, Valentin Emiya, Cédric Févotte
    Convex nonnegative matrix factorization with missing data Ronan Hamon, Valentin Emiya, Cédric Févotte To cite this version: Ronan Hamon, Valentin Emiya, Cédric Févotte. Convex nonnegative matrix factorization with missing data. IEEE International Workshop on Machine Learning for Signal Processing, Sep 2016, Vietri sul Mare, Salerno, Italy. hal-01346492 HAL Id: hal-01346492 https://hal-amu.archives-ouvertes.fr/hal-01346492 Submitted on 7 Oct 2016 HAL is a multi-disciplinary open access L’archive ouverte pluridisciplinaire HAL, est archive for the deposit and dissemination of sci- destinée au dépôt et à la diffusion de documents entific research documents, whether they are pub- scientifiques de niveau recherche, publiés ou non, lished or not. The documents may come from émanant des établissements d’enseignement et de teaching and research institutions in France or recherche français ou étrangers, des laboratoires abroad, or from public or private research centers. publics ou privés. 2016 IEEE INTERNATIONAL WORKSHOP ON MACHINE LEARNING FOR SIGNAL PROCESSING, SEPT. 13–16, 2016, SALERNO, ITALY CONVEX NONNEGATIVE MATRIX FACTORIZATION WITH MISSING DATA Ronan Hamon, Valentin Emiya˚ Cedric´ Fevotte´ Aix Marseille Univ, CNRS, LIF, Marseille, France CNRS & IRIT, Toulouse, France ABSTRACT example concerns the field of image or audio inpainting [5, 6, 7, 8], where CNMF may improve the current reconstruction Convex nonnegative matrix factorization (CNMF) is a variant techniques. In inpainting of audio spectrograms for example, of nonnegative matrix factorization (NMF) in which the com- setting up the dictionary to be a comprehensive collection of ponents are a convex combination of atoms of a known dic- notes from a specific instrument may guide the factorization tionary.
    [Show full text]
  • From Predictive Methods to Missing Data Imputation: an Optimization Approach
    Journal of Machine Learning Research 18 (2018) 1-39 Submitted 2/17; Revised 3/18; Published 4/18 From Predictive Methods to Missing Data Imputation: An Optimization Approach Dimitris Bertsimas [email protected] Colin Pawlowski [email protected] Ying Daisy Zhuo [email protected] Sloan School of Management and Operations Research Center Massachusetts Institute of Technology Cambridge, MA 02139 Editor: Francis Bach Abstract Missing data is a common problem in real-world settings and for this reason has attracted significant attention in the statistical literature. We propose a flexible framework based on formal optimization to impute missing data with mixed continuous and categorical variables. This framework can readily incorporate various predictive models including K- nearest neighbors, support vector machines, and decision tree based methods, and can be adapted for multiple imputation. We derive fast first-order methods that obtain high quality solutions in seconds following a general imputation algorithm opt.impute presented in this paper. We demonstrate that our proposed method improves out-of-sample accuracy in large-scale computational experiments across a sample of 84 data sets taken from the UCI Machine Learning Repository. In all scenarios of missing at random mechanisms and various missing percentages, opt.impute produces the best overall imputation in most data sets benchmarked against five other methods: mean impute, K-nearest neighbors, iterative knn, Bayesian PCA, and predictive-mean matching, with an average reduction in mean absolute error of 8.3% against the best cross-validated benchmark method. Moreover, opt.impute leads to improved out-of-sample performance of learning algorithms trained using the imputed data, demonstrated by computational experiments on 10 downstream tasks.
    [Show full text]