machine learning & knowledge extraction Review Introduction to Survival Analysis in Practice Frank Emmert-Streib 1,2,∗ and Matthias Dehmer 3,4,5 1 Predictive Society and Data Analytics Lab, Faculty of Information Technolgy and Communication Sciences, Tampere University, FI-33101 Tampere, Finland 2 Institute of Biosciences and Medical Technology, FI-33101 Tampere, Finland 3 Steyr School of Management, University of Applied Sciences Upper Austria, 4400 Steyr Campus, Austria 4 Department of Biomedical Computer Science and Mechatronics, UMIT- The Health and Life Science University, 6060 Hall in Tyrol, Austria 5 College of Artificial Intelligence, Nankai University, Tianjin 300350, China * Correspondence:
[email protected]; Tel.: +358-50-301-5353 Received: 31 July 2019; Accepted: 2 September 2019; Published: 8 September 2019 Abstract: The modeling of time to event data is an important topic with many applications in diverse areas. The collective of methods to analyze such data are called survival analysis, event history analysis or duration analysis. Survival analysis is widely applicable because the definition of an ’event’ can be manifold and examples include death, graduation, purchase or bankruptcy. Hence, application areas range from medicine and sociology to marketing and economics. In this paper, we review the theoretical basics of survival analysis including estimators for survival and hazard functions. We discuss the Cox Proportional Hazard Model in detail and also approaches for testing the proportional hazard (PH) assumption. Furthermore, we discuss stratified Cox models for cases when the PH assumption does not hold. Our discussion is complemented with a worked example using the statistical programming language R to enable the practical application of the methodology.