Thursday, 4 October 2012

Hadoop online training - Index


Table of Contents

Welcome to the Hadoop Online Training Tutorial. This tutorial includes the following materials designed to teach you how to use the Hadoop distributed data processing environment:
  • Hadoop 0.18.0 distribution (includes full source code)
  • A virtual machine image running Ubuntu Linux and preconfigured with Hadoop
  • VMware Player software to run the virtual machine image
  • A tutorial which will guide you through many aspects of Hadoop's installation and operation.
The tutorial is divided into seven modules, designed to be worked through in order. They can be accessed from the links below.
  1. Tutorial Introduction
  2. The Hadoop Distributed File System
  3. Getting Started With Hadoop
  4. MapReduce
  5. Advanced MapReduce Features
  6. Related Topics
  7. Managing a Hadoop Cluster
  8. Pig Tutorial
.

Hadoop Admin Course Content Details


Introduction to Big Data
• Characteristics of Big Data
• Why is parallel computing important
• Discuss various products developed by vendors
Introducing Hadoop
• Components of Hadoop
• Starting Hadoop
• Identify various processes
• Hands on
Working with HDFS
• Basic file commands
• Web Based User Interface
• Reading & Writing to files
• Run a word count program
• View jobs in the Web UI
• Hands on
Installation & Configuration of Hadoop
• Types of installation (RPM’s & Tar files)
• Set up ‘ssh’ for the Hadoop cluster
• Tree structure
• XML, masters and slaves files
• Checking system health
• Discuss block size and replication factor
• Benchmarking the cluster
• Hands on
Advanced administration activities
• Adding and de-commissioning nodes
• Purpose of secondary name node
• Recovery from a failed name node
• Managing quotas
• Enabling trash
• Hands on
Monitoring the Hadoop Cluster
• Hadoop infrastructure monitoring
• Hadoop specific monitoring
• Install and configure Nagios / Ganglia
• Capture metrics
• Hands on
Other Components of the Hadoop ecosystem
• Discuss Hive, Sqoop, Pig, HBase, Flume
• Use cases of each
• Use Hadoop streaming to write code in Perl / Python
• Hands on


Hadoop Developer Course Content Details


Introduction
The Motivation for Hadoop
·         Problems with traditional large-scale systems
·         Requirements for a new approach
Hadoop: Basic Concepts
·         An Overview of Hadoop
·         The Hadoop Distributed File System
·         Hands-On Exercise
·         How MapReduce Works
·         Hands-On Exercise
·         Anatomy of a Hadoop Cluster
·         Other Hadoop Ecosystem Components
Writing a MapReduce Program
·         The MapReduce Flow
·         Examining a Sample MapReduce Program
·         Basic MapReduce API Concepts
·         The Driver Code
·         The Mapper
·         The Reducer
·         Hadoop’s Streaming API
·         Using Eclipse for Rapid Development
·         Hands-on exercise
·         The New MapReduce API
Delving Deeper Into The Hadoop API
·         More about ToolRunner
·         Testing with MRUnit
·         Reducing Intermediate Data With Combiners
·         The configure and close methods for Map/Reduce Setup and Teardown
·         Writing Partitioners for Better Load Balancing
·         Hands-On Exercise
·         Directly Accessing HDFS
·         Using the Distributed Cache
·         Hands-On Exercise
Common MapReduce Algorithms
·         Sorting and Searching
·         Indexing
·         Machine Learning With Mahout
·         Term Frequency – Inverse Document Frequency
·         Word Co-Occurrence
·         Hands-On Exercise
Usining HBase
·         What is HBase?
·         HBase Architecture
·         HBase API
·         Managing large data sets with HBase
·         Using HBase in Hadoop applications
·         Hands-on exercise
Using Hive and Pig
·         Hive Basics
·         Pig Basics
·         Hands-on exercise
Practical Development Tips and Techniques
·         Debugging MapReduce Code
·         Using LocalJobRunner Mode For Easier Debugging
·         Retrieving Job Information with Counters
·         Logging
·         Splittable File Formats
·         Determining the Optimal Number of Reducers
·         Map-Only MapReduce Jobs
·         Hands-On Exercise

More Advanced MapReduce Programming
·         Custom Writables and WritableComparables
·         Saving Binary Data using SequenceFiles and Avro Files
·         Creating InputFormats and OutputFormats
·         Hands-On Exercise
Joining Data Sets in MapReduce
·         Map-Side Joins
·         The Secondary Sort
·         Reduce-Side Joins