BIG DATA
Learn with HADWIK New Technologies Shaping Today's Big Data World
Learn with HADWIK New Technologies Shaping Today's Big Data World
Understand the Big Data Platform and its Use cases
Provide an overview of Apache Hadoop
Provide HDFS Concepts and Interfacing with HDFS
Understand Map Reduce Jobs
Provide hands on Hadoop Eco System
Apply analytics on Structured, Unstructured Data.
Exposure to Data Analytics with R
What you will Learn !
Module 1 : INTRODUCTION TO BIG DATA AND HADOOP
Types of Digital Data, Introduction to Big Data, Big Data Analytics, History of Hadoop, Apache Hadoop, Analysing Data with Unix tools, Analysing Data with Hadoop, Hadoop Streaming, Hadoop Echo System, IBM Big Data Strategy, Introduction to Infosphere Big Insights and Big Sheets.
Module 2 : HDFS(Hadoop Distributed File System)
The Design of HDFS, HDFS Concepts, Command Line Interface, Hadoop file system interfaces, Data flow, Data Ingest with Flume and Scoop and Hadoop archives, Hadoop I/O: Compression, Serialization, Avro and File-Based Data structures.
Module 3 : Map Reduce
Anatomy of a Map Reduce Job Run, Failures, Job Scheduling, Shuffle and Sort, Task Execution, Map Reduce Types and Formats, Map Reduce Features.
Module 4 : Hadoop Eco System Pig :
Pig : Introduction to PIG, Execution Modes of Pig, Comparison of Pig with Databases, Grunt, Pig Latin, User Defined Functions, Data Processing operators.
Hive : Hive Shell, Hive Services, Hive Metastore, Comparison with Traditional Databases, HiveQL, Tables, Querying Data and User Defined Functions.
Hbase : HBasics, Concepts, Clients, Example, Hbase Versus RDBMS.
Big SQL : Introduction
Module 5 : Data Analytics with R Machine Learning :
Machine Learning : Introduction, Supervised Learning, Unsupervised Learning, Collaborative Filtering. Big Data Analytics with BigR.
Requirement :
Should have knowledge of one Programming Language (Java preferably), Practice of SQL (queries and sub queries), exposure to Linux Environment.
Duration
60 Hrs.
Project
1
Support
Lifetime
Certification
Yes