In-Memory Caching Orchestration for Hadoop

descriptionPublicationkeyboard_double_arrow_right Article , Conference object 01 May 2016Publisher:IEEEJournal:2016 16th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid)

Authors: Jaewon Kwak; Eunji Hwang; Tae-kyung Yoo; Beomseok Nam; Young-ri Choi;

doi: 10.1109/ccgrid.2016.73

In-Memory Caching Orchestration for Hadoop

- Summary
- Metrics

Abstract

In this paper, we investigate techniques to effectively orchestrate HDFS in-memory caching for Hadoop. We first evaluate a degree of benefit which each of various MapReduce applications can get from in-memory caching, i.e. cache affinity. We then propose an adaptive cache local scheduling algorithm that adaptively adjusts the waiting time of a MapReduce job in a queue for a cache local node. We set the waiting time to be proportional to the percentage of cached input data for the job. We also develop a cache affinity cache replacement algorithm that determines which block is cached and evicted based on the cache affinity of applications. Using various workloads consisting of multiple MapReduce applications, we conduct experimental study to demonstrate the effects of the proposed in-memory orchestration techniques. Our experimental results show that our enhanced Hadoop in-memory caching scheme improves the performance of the MapReduce workloads up to 18% and 10% against Hadoop that disables and enables HDFS in-memory caching, respectively.

Related Organizations

Ulsan National Institute of Science and Technology
Korea (Republic of)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	8
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%