On a Dynamic Data Placement Strategy for Heterogeneous Hadoop Clusters

Yang Liu, Chase Q. Wu, Meng Wang, Aiqin Hou, Yongqiang Wang

Research output: Chapter in Book/Report/Conference proceedingConference contribution

12 Scopus citations

Abstract

Hadoop is one of the most popular distributed systems for big data computing in both industry and science communities. The default data placement strategy of Hadoop Distributed File System (HDFS), which was initially designed for homogenous environments, may suffer from performance degradation when deployed in heterogeneous clusters comprised of data nodes with disparate computing power and disk capacity, hence undermining the performance of MapReduce applications. In this paper, we use a Grey Forecast model to predict data hotness dynamically and determine an appropriate number of data block replicas on the fly. Based on such information, we further propose a dynamic data placement strategy (DDPS) to decide the best location for new replicas according to their hotness. The proposed method is able to dynamically adjust data replicas stored on each node in a heterogeneous Hadoop cluster and reduce the response time of big data applications. Experimental results on a heterogeneous Hadoop cluster show that DDPS together with the prediction model significantly increases application execution efficiency and improve MapReduce performance over the default HDFS configuration.

Original languageEnglish (US)
Title of host publication2018 International Symposium on Networks, Computers and Communications, ISNCC 2018
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9781538637784
DOIs
StatePublished - Nov 9 2018
Event2018 International Symposium on Networks, Computers and Communications, ISNCC 2018 - Rome, Italy
Duration: Jun 19 2018Jun 21 2018

Publication series

Name2018 International Symposium on Networks, Computers and Communications, ISNCC 2018

Other

Other2018 International Symposium on Networks, Computers and Communications, ISNCC 2018
Country/TerritoryItaly
CityRome
Period6/19/186/21/18

All Science Journal Classification (ASJC) codes

  • Computational Theory and Mathematics
  • Computer Networks and Communications
  • Hardware and Architecture
  • Energy Engineering and Power Technology
  • Safety, Risk, Reliability and Quality

Keywords

  • HDFS
  • Heterogeneous Hadoop
  • MapReduce
  • data placement
  • prediction model

Fingerprint

Dive into the research topics of 'On a Dynamic Data Placement Strategy for Heterogeneous Hadoop Clusters'. Together they form a unique fingerprint.

Cite this