Hadoop 2 Introduces YARN
Hadoop 2 Introduces YARN
Written by Kay Ewbank   
Tuesday, 22 October 2013

Apache Hadoop 2 is now available with support for Apache YARN, a framework for job scheduling and cluster resource management, and high availability for the HDFS filing system.

In fact, the release is 2.2.0, but it is the first stable release in the 2.x line. The biggest change in the new version of the distributed big data storage and analysis framework is the support for YARN. YARN sits on top of HDFS (Hadoop distributed file system) and serves as a large-scale, distributed operating system for big data applications, so multiple applications can run simultaneously.

This is a major change to the way Hadoop works. Previous releases relied on the MapReduce framework to handle the division of work as well as managing the resources of the servers. In the new version YARN is used as the resource manager.

 

Architectural View of YARN (Source: HortonWorks)

The way YARN works is by splitting the work undertaken by the MapReduce JobTracker component in two. JobTracker has until now managed both resource management and job scheduling/monitoring, but these are now run as two separate applications: a global ResourceManager and an ApplicationMaster that has a separate copy for each application running.

If you want to find out more about YARN, Arun Murthy, release manager of Apache Hadoop 2 and Founder of Hortonworks, outlines it in an informative series of posts on the Hortonworks blog.

This change opens up Hadoop so that developers will be able to build apps directly within Hadoop, instead of writing them to run externally, and represents a major opening up of Hadoop.

Alongside the introduction of YARN, the HDFS aspect of Hadoop has also been improved, with high availability for HDFS, support for HDFS snapshots so you’ll be able to use native backup and disaster recovery processes. The ability to use the NFSv3 file system to access data in HDFS also makes Hadoop more ‘mainstream’, as it can now be mounted as a standard filesystem. Native network encryption has been added to secure data in transit. The team behind Hadoop has put a lot of work into stabilizing Hadoop’s APIs. Finally, Hadoop 2.2 also adds support for running it on Microsoft Windows.

 

hadoopsquare

More Information
Apache Hadoop

Introducing Apache Hadoop YARN

Related Articles

Hadoop for Windows

Microsoft Open Sources Big Data REEF

Hadoop gets to 1.0

Hive on Hadoop for MongoDB

 

To be informed about new articles on I Programmer, install the I Programmer Toolbar, subscribe to the RSS feed, follow us on, Twitter, Facebook, Google+ or Linkedin,  or sign up for our weekly newsletter.

 

 
 
 

blog comments powered by Disqus

 

Banner


Arduino 1.8.0 IDE For Reunified Arduino
29/12/2016

This new version of the Arduino IDE was announced at the beginning of October 2016 at the New York Maker Faire as the official, unified, desktop editor for all Arduino boards.



Big Increase in AI, Cognitive and Cloud Computing Patents in 2016
16/01/2017

IBM and Samsung head the charts for most US Patent grants in 2016. Which comes top depends on whose statistics you choose, but more interesting is what the patents relate to.


More News

 

Last Updated ( Tuesday, 22 October 2013 )
 
 

   
RSS feed of news items only
I Programmer News
Copyright © 2017 i-programmer.info. All Rights Reserved.
Joomla! is Free Software released under the GNU/GPL License.