Professional Summary Technical Skills
Total Page:16
File Type:pdf, Size:1020Kb
Load more
Recommended publications
-
Use Splunk with Big Data Repositories Like Spark, Solr, Hadoop and Nosql Storage
Copyright © 2016 Splunk Inc. Use Splunk With Big Data Repositories Like Spark, Solr, Hadoop And Nosql Storage Raanan Dagan, May Long Big Data Architect, Splunk Disclaimer During the course of this presentaon, we may make forward looking statements regarding future events or the expected performance of the company. We cauJon you that such statements reflect our current expectaons and esJmates based on factors currently known to us and that actual events or results could differ materially. For important factors that may cause actual results to differ from those contained in our forward-looking statements, please review our filings with the SEC. The forward- looking statements made in the this presentaon are being made as of the Jme and date of its live presentaon. If reviewed aer its live presentaon, this presentaon may not contain current or accurate informaon. We do not assume any obligaon to update any forward looking statements we may make. In addiJon, any informaon about our roadmap outlines our general product direcJon and is subject to change at any Jme without noJce. It is for informaonal purposes only and shall not, be incorporated into any contract or other commitment. Splunk undertakes no obligaon either to develop the features or funcJonality described or to include any such feature or funcJonality in a future release. 2 Agenda Use Cases: Fraud With Solr, Splunk, And Splunk AnalyJcs For Hadoop Business AnalyJcs With Cassandra, Splunk Cloud, And Splunk AnalyJcs For Hadoop Document Classificaon With Spark And Splunk Network IT With Kaa And Splunk Kaa Add On Demo 3 Fraud With Solr, Splunk, And Splunk AnalyJcs For Hadoop Use Case: Fraud – Why Apache Solr Apache Solr is an open source enterprise search plaorm from the Apache Lucene API. -
Personal Information Backgrounds Abilities Highlights
Yi Du Computer Network Information Center, Chinese Academy of Sciences Phone:+86-15810134970 Email:[email protected] Homepage: http://yiducn.github.io/ Personal Information Gender: Male Date of Graduate:July, 2013 Address: 4# Building, South 4th Street Zhong Guan Cun. Beijing, P.R.China. 100190 Backgrounds 2015.12~Now Department of Big Data Technology and Application Development, Computer Network Information Center Beijing, China Job Title: Associate Professor Research Interest: Data Mining, Visual Analytics 2015.09~2016.09 School of Electrical and Computer Engineering, Purdue University USA Job Title: Visiting Scholar Research Interest: Spatio-temporal Visualization, Visual Analytics 2013.09~2015.12 Scientific Data Center, Computer Network Information Center, CAS Beijing, China Job Title: Assistant Professor Research Interest: Data Processing, Data Visualization, HCI 2008.09~2013.09 Institute of Software Chinese Academy of Sciences Beijing, China Major: Computer Applied Technology Doctoral Degree Research Interest: Human Computer Interaction(HCI), Information Visualization 2004.08-2008.07 Shandong University Jinan Major: Software Engineering Bachelor's Degree Abilities Master the design and development of data science system, including data collecting, wrangling, analyzing, mining, visualizing and interacting. Master analyzing and mining of large-scale spatio-temporal data. Experiences in coding with Java, JavaScript. Familiar with Python and C++. Experiences in MongoDB, DB2. Familiar with MS SQLServer and Oracle. Experiences in traditional data mining and machine learning algorithms. Familiar with Hadoop, Titan, Kylin, etc. Highlights Participated in an open source project Gephi, contributed several plugins with over 3000 lines of code. An analytics platform named DVIZ, in which I played the leading role, gained several prizes in China. -
Hortonworks Cybersecurity Platform Administration (April 24, 2018)
Hortonworks Cybersecurity Platform Administration (April 24, 2018) docs.cloudera.com Hortonworks Cybersecurity April 24, 2018 Platform Hortonworks Cybersecurity Platform: Administration Copyright © 2012-2018 Hortonworks, Inc. Some rights reserved. Hortonworks Cybersecurity Platform (HCP) is a modern data application based on Apache Metron, powered by Apache Hadoop, Apache Storm, and related technologies. HCP provides a framework and tools to enable greater efficiency in Security Operation Centers (SOCs) along with better and faster threat detection in real-time at massive scale. It provides ingestion, parsing and normalization of fully enriched, contextualized data, threat intelligence feeds, triage and machine learning based detection. It also provides end user near real-time dashboarding. Based on a strong foundation in the Hortonworks Data Platform (HDP) and Hortonworks DataFlow (HDF) stacks, HCP provides an integrated advanced platform for security analytics. Please visit the Hortonworks Data Platform page for more information on Hortonworks technology. For more information on Hortonworks services, please visit either the Support or Training page. Feel free to Contact Us directly to discuss your specific needs. Except where otherwise noted, this document is licensed under Creative Commons Attribution ShareAlike 4.0 License. http://creativecommons.org/licenses/by-sa/4.0/legalcode ii Hortonworks Cybersecurity April 24, 2018 Platform Table of Contents 1. HCP Information Roadmap ......................................................................................... -
My Steps to Learn About Apache Nifi
My steps to learn about Apache NiFi Paulo Jerônimo, 2018-05-24 05:36:18 WEST Table of Contents Introduction. 1 About this document . 1 About me . 1 Videos with a technical background . 2 Lab 1: Running Apache NiFi inside a Docker container . 3 Prerequisites . 3 Start/Restart. 3 Access to the UI . 3 Status. 3 Stop . 3 Lab 2: Running Apache NiFi locally . 5 Prerequisites . 5 Installation. 5 Start . 5 Access to the UI . 5 Status. 5 Stop . 6 Lab 3: Building a simple Data Flow . 7 Prerequisites . 7 Step 1 - Create a Nifi docker container with default parameters . 7 Step 2 - Access the UI and create two processors . 7 Step 3 - Add and configure processor 1 (GenerateFlowFile) . 7 Step 4 - Add and configure processor 2 (Putfile) . 10 Step 5 - Connect the processors . 12 Step 6 - Start the processors. 14 Step 7 - View the generated logs . 14 Step 8 - Stop the processors . 15 Step 9 - Stop and destroy the docker container . 15 Conclusions . 15 All references . 16 Introduction Recently I had work to produce a document with a comparison between two tools for Cloud Data Flow. I didn’t have any knowledge of this kind of technology before creating this document. Apache NiFi is one of the tools in my comparison document. So, here I describe some of my procedures to learn about it and take my own preliminary conclusions. I followed many steps on my own desktop (a MacBook Pro computer) to accomplish this task. This document shows you what I did. Basically, to learn about Apache NiFi in order to do a comparison with other tool: • I saw some videos about it. -
Extended Version
Sina Sheikholeslami C u rriculum V it a e ( Last U pdated N ovember 2 0 18) Website: http://sinash.ir Present Address : https://www.kth.se/profile/sinash EIT Digital Stockholm CLC , https://linkedin.com/in/sinasheikholeslami Isafjordsgatan 26, Email: si [email protected] 164 40 Kista (Stockholm), [email protected] Sweden [email protected] Educational Background: • M.Sc. Student of Data Science, Eindhoven University of Technology & KTH Royal Institute of Technology, Under EIT-Digital Master School. 2017-Present. • B.Sc. in Computer Software Engineering, Department of Computer Engineering and Information Technology, Amirkabir University of Technology (Tehran Polytechnic). 2011-2016. • Mirza Koochak Khan Pre-College in Mathematics and Physics, Rasht, National Organization for Development of Exceptional Talents (NODET). Overall GPA: 19.61/20. 2010-2011. • Mirza Koochak Khan Highschool in Mathematics and Physics, Rasht, National Organization for Development of Exceptional Talents (NODET). Overall GPA: 19.17/20, Final Year's GPA: 19.66/20. 2007-2010. Research Fields of Interest: • Distributed Deep Learning, Hyperparameter Optimization, AutoML, Data Intensive Computing Bachelor's Thesis: • “SDMiner: A Tool for Mining Data Streams on Top of Apache Spark”, Under supervision of Dr. Amir H. Payberah (Defended on June 29th 2016). Computer Skills: • Programming Languages & Markups: o F luent in Java, Python, Scala, JavaS cript, C/C++, A ndroid Pr ogram Develop ment o Familia r wit h R, SAS, SQL , Nod e.js, An gula rJS, HTM L, JSP • -
Handling Data Flows of Streaming Internet of Things Data
IT16048 Examensarbete 30 hp Juni 2016 Handling Data Flows of Streaming Internet of Things Data Yonatan Kebede Serbessa Masterprogram i datavetenskap Master Programme in Computer Science i Abstract Handling Data Flows of Streaming Internet of Things Data Yonatan Kebede Serbessa Teknisk- naturvetenskaplig fakultet UTH-enheten Streaming data in various formats is generated in a very fast way and these data needs to be processed and analyzed before it becomes useless. The technology currently Besöksadress: existing provides the tools to process these data and gain more meaningful Ångströmlaboratoriet Lägerhyddsvägen 1 information out of it. This thesis has two parts: theoretical and practical. The Hus 4, Plan 0 theoretical part investigates what tools are there that are suitable for stream data flow processing and analysis. In doing so, it starts with studying one of the main Postadress: streaming data source that produce large volumes of data: Internet of Things. In this, Box 536 751 21 Uppsala the technologies behind it, common use cases, challenges, and solutions are studied. Then it is followed by overview of selected tools namely Apache NiFi, Apache Spark Telefon: Streaming and Apache Storm studying their key features, main components, and 018 – 471 30 03 architecture. After the tools are studied, 5 parameters are selected to review how Telefax: each tool handles these parameters. This can be useful for considering choosing 018 – 471 30 00 certain tool given the parameters and the use case at hand. The second part of the thesis involves Twitter data analysis which is done using Apache NiFi, one of the tools Hemsida: studied. The purpose is to show how NiFi can be used for processing data starting http://www.teknat.uu.se/student from ingestion to finally sending it to storage systems. -
Poweredge R640 Apache Hadoop
A Principled Technologies report: Hands-on testing. Real-world results. The science behind the report: Run compute-intensive Apache Hadoop big data workloads faster with Dell EMC PowerEdge R640 servers This document describes what we tested, how we tested, and what we found. To learn how these facts translate into real-world benefits, read the report Run compute-intensive Apache Hadoop big data workloads faster with Dell EMC PowerEdge R640 servers. We concluded our hands-on testing on October 27, 2019. During testing, we determined the appropriate hardware and software configurations and applied updates as they became available. The results in this report reflect configurations that we finalized on October 15, 2019 or earlier. Unavoidably, these configurations may not represent the latest versions available when this report appears. Our results The table below presents the throughput each solution delivered when running the HiBench workloads. Dell EMC™ PowerEdge™ R640 Dell EMC PowerEdge R630 Percentage more throughput solution solution Latent Dirichlet Allocation 4.13 1.94 112% (MB/sec) Random Forest (MB/sec) 100.66 94.43 6% WordCount (GB/sec) 5.10 3.45 47% The table below presents the minutes each solution needed to complete the HiBench workloads. Dell EMC PowerEdge R640 Dell EMC PowerEdge R630 Percentage less time solution solution Latent Dirichlet Allocation 17.11 36.25 52% Random Forest 5.55 5.92 6% WordCount 4.95 7.32 32% Run compute-intensive Apache Hadoop big data workloads faster with Dell EMC PowerEdge R640 servers November 2019 System configuration information The table below presents detailed information on the systems we tested. -
TR-4744: Secure Hadoop Using Apache Ranger with Netapp In
Technical Report Secure Hadoop using Apache Ranger with NetApp In-Place Analytics Module Deployment Guide Karthikeyan Nagalingam, NetApp February 2019 | TR-4744 Abstract This document introduces the NetApp® In-Place Analytics Module for Apache Hadoop and Spark with Ranger. The topics covered in this report include the Ranger configuration, underlying architecture, integration with Hadoop, and benefits of Ranger with NetApp In-Place Analytics Module using Hadoop with NetApp ONTAP® data management software. TABLE OF CONTENTS 1 Introduction ........................................................................................................................................... 4 1.1 Overview .........................................................................................................................................................4 1.2 Deployment Options .......................................................................................................................................5 1.3 NetApp In-Place Analytics Module 3.0.1 Features ..........................................................................................5 2 Ranger ................................................................................................................................................... 6 2.1 Components Validated with Ranger ................................................................................................................6 3 NetApp In-Place Analytics Module Design with Ranger.................................................................. -
Hortonworks Data Platform
Hortonworks Data Platform Apache Ambari Installation for IBM Power Systems (November 15, 2018) docs.cloudera.com Hortonworks Data Platform November 15, 2018 Hortonworks Data Platform: Apache Ambari Installation for IBM Power Systems Copyright © 2012-2018 Hortonworks, Inc. Some rights reserved. The Hortonworks Data Platform, powered by Apache Hadoop, is a massively scalable and 100% open source platform for storing, processing and analyzing large volumes of data. It is designed to deal with data from many sources and formats in a very quick, easy and cost-effective manner. The Hortonworks Data Platform consists of the essential set of Apache Hadoop projects including MapReduce, Hadoop Distributed File System (HDFS), HCatalog, Pig, Hive, HBase, ZooKeeper and Ambari. Hortonworks is the major contributor of code and patches to many of these projects. These projects have been integrated and tested as part of the Hortonworks Data Platform release process and installation and configuration tools have also been included. Unlike other providers of platforms built using Apache Hadoop, Hortonworks contributes 100% of our code back to the Apache Software Foundation. The Hortonworks Data Platform is Apache-licensed and completely open source. We sell only expert technical support, training and partner-enablement services. All of our technology is, and will remain free and open source. Please visit the Hortonworks Data Platform page for more information on Hortonworks technology. For more information on Hortonworks services, please visit either the Support or Training page. Feel free to Contact Us directly to discuss your specific needs. Except where otherwise noted, this document is licensed under Creative Commons Attribution ShareAlike 4.0 License. -
Hdf® Stream Developer 3 Days
TRAINING OFFERING | DEV-371 HDF® STREAM DEVELOPER 3 DAYS This course is designed for Data Engineers, Data Stewards and Data Flow Managers who need to automate the flow of data between systems as well as create real-time applications to ingest and process streaming data sources using Hortonworks Data Flow (HDF) environments. Specific technologies covered include: Apache NiFi, Apache Kafka and Apache Storm. The course will culminate in the creation of a end-to-end exercise that spans this HDF technology stack. PREREQUISITES Students should be familiar with programming principles and have previous experience in software development. First-hand experience with Java programming and developing within an IDE are required. Experience with Linux and a basic understanding of DataFlow tools and would be helpful. No prior Hadoop experience required. TARGET AUDIENCE Developers, Data & Integration Engineers, and Architects who need to automate data flow between systems and/or develop streaming applications. FORMAT 50% Lecture/Discussion 50% Hands-on Labs AGENDA SUMMARY Day 1: Introduction to HDF Components, Apache NiFi dataflow development Day 2: Apache Kafka, NiFi integration with HDF/HDP, Apache Storm architecture Day 3: Storm management options, multi-language support, Kafka integration DAY 1 OBJECTIVES • Introduce HDF’s components; Apache NiFi, Apache Kafka, and Apache Storm • NiFi architecture, features, and characteristics • NiFi user interface; processors and connections in detail • NiFi dataflow assembly • Processor Groups and their elements -
Hortonworks Dataflow Getting Started (January 31, 2018)
Hortonworks DataFlow Getting Started (January 31, 2018) docs.cloudera.com Hortonworks DataFlow January 31, 2018 Hortonworks DataFlow: Getting Started Copyright © 2012-2018 Hortonworks, Inc. Some rights reserved. Except where otherwise noted, this document is licensed under Creative Commons Attribution ShareAlike 4.0 License. http://creativecommons.org/licenses/by-sa/4.0/legalcode ii Hortonworks DataFlow January 31, 2018 Table of Contents 1. Getting Started with Apache NiFi ................................................................................ 1 1.1. Getting Started with Apache NiFi ...................................................................... 1 1.1.1. Terminology Used in This Guide ............................................................. 1 1.1.2. Starting NiFi ........................................................................................... 1 1.1.3. I Started NiFi. Now What? ..................................................................... 2 1.1.4. What Processors are Available ................................................................ 7 1.1.5. Working With Attributes ...................................................................... 13 1.1.6. Custom Properties Within Expression Language .................................... 16 1.1.7. Working With Templates ...................................................................... 17 1.1.8. Monitoring NiFi .................................................................................... 18 1.1.9. Data Provenance ................................................................................. -
End to End Data Flow Management and Streaming Analytics Platform
SOLUTION SHEET End to End Data Flow Management and Streaming Analytics Platform CREATE STREAMING ANALYTICS APPLICATIONS IN MINUTES WITHOUT WRITING CODE The increasing growth of data, especially data-in-motion, presents Immediate and Continuous Insights enterprises with the challenges of managing streaming data and How do you analyze data-in-motion when it has not landed in a getting actionable intelligence. Hortonworks DataFlow (HDF™) database yet? HDF’s Streaming Analytics Manager (SAM) feature provides the only end-to-end streaming data platform with flow allows organizations to create analytics applications in minutes management, stream processing, and enterprise services which to capture perishable insights in real-time without writing a collect, curate, analyze and act on data in the data center and single line of code. Streaming Analytics Manager is a tool used cloud. Complementary to the Hortonworks Data Platform (HDP®), to design, develop, deploy and manage streaming analytics HDF is powered by key open sourced projects including Apache applications using a drag-drop visual paradigm. A developer can NiFi/MiniFi, Apache Kafka®, Apache Storm™, and Druid. build complex streaming analytics applications without having to know the complexities of the underlying streaming engine. • Easy, Flexible, Secure Way to Get the Data You Need The biggest challenge to getting data insights to work • Enterprise Grade Corporate Governance, Security and for your organization is getting the data in the first place; Operations ingestion, cleansing, and preparing the data for analysis. This Streaming data needs to meet the same enterprise corporate is complicated by data-in-motion, which could operate under governance and security standards for operations as varying conditions such as velocity and bandwidth over a other traditional data types.