Bayesian Networks for Root Cause Analysis

Model driven root cause analysis and reliability enhancement for large distributed computing systems Michał Zasadziński ADVERTIMENT La consulta d’aquesta tesi queda condicionada a l’acceptació de les següents condicions d'ús: La difusió d’aquesta tesi per mitjà del repositori institucional UPCommons (http://upcommons.upc.edu/tesis) i el repositori cooperatiu TDX ( http://www.tdx.cat/ ) ha estat autoritzada pels titulars dels drets de propietat intel·lectual únicament per a usos privats emmarcats en activitats d’investigació i docència. No s’autoritza la seva reproducció amb finalitats de lucre ni la seva difusió i posada a disposició des d’un lloc aliè al servei UPCommons o TDX. No s’autoritza la presentació del seu contingut en una finestra o marc aliè a UPCommons (framing). Aquesta reserva de drets afecta tant al resum de presentació de la tesi com als seus continguts. En la utilització o cita de parts de la tesi és obligat indicar el nom de la persona autora. ADVERTENCIA La consulta de esta tesis queda condicionada a la aceptación de las siguientes condiciones de uso: La difusión de esta tesis por medio del repositorio institucional UPCommons (http://upcommons.upc.edu/tesis) y el repositorio cooperativo TDR (http://www.tdx.cat/?locale- attribute=es) ha sido autorizada por los titulares de los derechos de propiedad intelectual únicamente para usos privados enmarcados en actividades de investigación y docencia. No se autoriza su reproducción con finalidades de lucro ni su difusión y puesta a disposición desde un sitio ajeno al servicio UPCommons No se autoriza la presentación de su contenido en una ventana o marco ajeno a UPCommons (framing). Esta reserva de derechos afecta tanto al resumen de presentación de la tesis como a sus contenidos. En la utilización o cita de partes de la tesis es obligado indicar el nombre de la persona autora. WARNING On having consulted this thesis you’re accepting the following use conditions: Spreading this thesis by the institutional repository UPCommons (http://upcommons.upc.edu/tesis) and the cooperative repository TDX (http://www.tdx.cat/?locale- attribute=en) has been authorized by the titular of the intellectual property rights only for private uses placed in investigation and teaching activities. Reproduction with lucrative aims is not authorized neither its spreading nor availability from a site foreign to the UPCommons service. Introducing its content in a window or frame foreign to the UPCommons service is not authorized (framing). These rights affect to the presentation summary of the thesis as well as to its contents. In the using or citation of parts of the thesis it’s obliged to indicate the name of the author. DOCTORAL THESIS Model Driven Root Cause Analysis and Reliability Enhancement for Large Distributed Computing Systems M.Eng. Micha lZasadziński Computer Architecture Department Universitat Politècnicade Catalunya Barcelona, Spain Supervisors: Dr. David Carrera Dr. Victor Muntés-Mulero Thesis presented to obtain the qualification of Doctor from the Universitat Politècnicade Catalunya. October, 2018, Barcelona ii ACKNOWLEDGMENTS I would like to thank many extraordinary and fantastic people for their input to the work of this thesis. Firstly, thanks to my wife, parents and the rest of my family. It could have been impossible to accomplish this mission without their support, understanding and fresh views. I would like to thank Victor Muntés-Muleroand Marc Soléfor the mentorship and guidance they gave me. I could have observed what is called the right mentorship: \mentor stretches you, pushes for the risk, but at the same time shows the excellence and when you fail - support". It has been a pleasure to work with you. This PhD was intense and demanding, it had various phases, but the last one is exactly what I wish all young scholars and engineers: bravery and curiosity. Explore the World! Michal iii iv ABSTRACT Over last years the number of Big Data, supercomputing, Internet of Things or edge systems has snowballed. The core part of many areas and services in academia and business, are large, distributed, and complex Information Technology (IT) systems. Any failure or performance degradation occurring in these systems causes negative effects. The user experience is decreased, operational costs raises, and a system loses its availability and readiness for demanding service. Also, there is a higher environmental footprint, i.e., energy consumption. As the response, IT operators take care of resolving failures, issues, and unexpected events. Operators are aided with IT tools for monitoring, diagnostics, and root cause analysis. Usually, they have to troubleshoot and diagnose a system. Then, they perform some action to recover a system to its normal state. However, the characteristics of the emerging and future IT systems makes the diagnostics and root cause analysis hard and complicated. These systems can contain even billions of elements, distributed all over the Earth. Moreover, systems can frequently change, consequently making existing diagnostic models outdated for the effective operation. Even the most skillful operators have problems to deal with these systems to satisfy QoS constraints and deliver spectacular user experience. In this thesis, we contribute by making a step towards NoOps operating model. NoOps stands for no operations. In such a model of maintenance, an IT infrastructure is self-manageable and runs without human intervention. We would like to aid operator work and in the long term substitute them in the root cause analysis. We contribute for environments as mentioned earlier in two areas: (1) diagnostics, root cause analysis, root cause classification, and (2) prevention of failures. In particular, we focus on areas such as scalability, dynamism, lack of knowledge on system failures, predictability, and prevention of failures. For each of these aspects, we use different IT environment, for a proper diversification of the use cases. Firstly, we propose a fast root cause analysis (RCA) system based on probabilistic reasoning. The system can diagnose networks of devices with millions of nodes in a diagnostic model and solves the v problem of scalability of root cause analysis. In details, we leverage the fact that these systems usually contain repeatable elements. We create diagnostic models based on Bayesian networks. Then, we transform them into a more efficient structure for runtime use that are Arithmetic Circuits. Thanks to the proposed optimization in this transformation and cache-based mechanism, the solution performs better than state of the art techniques in terms of the memory consumption and time performance. Based on this contribution, we propose a root cause analysis system which works with dynamically changing environments. It is very common in Internet of Things environments, where devices connect and disconnect simultaneously, so the diagnostic model of the whole system changes rapidly. We propose actor based root cause analysis. This method is based on distributing diagnostic calculations through the devices and use of self-diagnostics paradigm. Thanks to this solution, results of partial diagnosis are known even when the connectivity with a part of the diagnosed system is lost. Also, we provide a framework to define supervising strategies which are utilized in case of a failure. We show that the contribution works well in a simulated Internet of Things system with high dynamism in its structure. Secondly, we focus on the aspect of knowledge integration and partial knowledge of a diagnosed system. The path to NoOps involves not only precise, reliable and fast diagnostics but also reusing as much knowledge as possible after the system is reconfigured or changed. The biggest challenge is to leverage knowledge on one IT system and reuse this knowledge for diagnostics of another, different system. We propose a weighted graph framework which can transfer knowledge and perform high- quality diagnostics of IT systems. We encode all possible data in a graph representation of a system state and automatically calculate weights of these graphs. Then, thanks to the similarity evaluation, we transfer knowledge about failures from one system to another and use it for diagnostics. We successfully evaluate the proposed approach on Spark, Hadoop, Kafka and Cassandra systems. For this purpose, we use an on-premise Big Data cluster and a cloud system of containers. Thirdly, we focus on the predictability of a supercomputing environment and prevention of failures. Failed jobs in a supercomputer cause not only waste in CPU time or energy consumption but also decrease work efficiency of users. Mining data collected during the operation of data centers helps to find patterns explaining failures and can be used to predict them. Automating system reactions, e.g., early termination of jobs, when software failures are predicted does not only increase availability and reduce operating cost, but it also frees administrators' and users' time. We explore a unique vi dataset containing the topology, operation metrics, and job scheduler history from the petascale Mistral supercomputer. We extract the most relevant system features deciding on the final state of a job through decision trees. Then, we successfully propose actions to prevent failures. We create a model to predict job evolution based on power time series of nodes. Finally, we evaluate the effect on CPU time saving for static and dynamic job termination policies. We finish the thesis with a short discussion on the contributions

Bayesian Networks for Root Cause Analysis

The Reliability Wall for Exascale Supercomputing Xuejun Yang, Member, IEEE, Zhiyuan Wang, Jingling Xue, Senior Member, IEEE, and Yun Zhou

LLNL Computation Directorate Annual Report (2014)

Biomolecular Simulation Data Management In

Measuring Power Consumption on IBM Blue Gene/P

Blue Gene®/Q Overview and Update

TMA4280—Introduction to Supercomputing

Introduction to the Practical Aspects of Programming for IBM Blue Gene/P

The Blue Gene/Q Compute Chip

R00456--FM Getting up to Speed

IBM Research Report Designing a Highly-Scalable Operating System

CFD) Applications on HPC Systems

2011 INCITE Awards Type: Renewal Title