Embedding GPU Computations in Hadoop

descriptionPublicationkeyboard_double_arrow_right Article 01 Jan 2014 English Publisher:Springer Science and Business Media LLCJournal:International Journal of Networked and Distributed Computing, volume 2, page 211 (eissn: 2211-7946,

Copyright policy )

Authors: Jie Zhu; Hai Jiang 0003; Juanjuan Li; Erikson Hardesty; Kuan-Ching Li; Zhongwen Li;

doi: 10.2991/ijndc.2014.2.4.2

Embedding GPU Computations in Hadoop

- Summary
- Subjects
- Metrics

Abstract

As the size of high performance applications increases, four major challenges including heterogeneity, programmability, fault resilience, and energy efficiency have arisen in the underlying distributed systems. To tackle with all of them without sacrificing performance, traditional approaches in resource utilization, task scheduling and programming paradigm should be reconsidered. While Hadoop has handled data-intensive applications well in Clouds, GPU has demonstrated its acceleration effectiveness for computation-intensive ones. This paper addresses the approaches for Hadoop to exploiting both CPU and GPU resources effectively to handle aforementioned challenges. Hadoop schedules MapReduce’s Map and Reduce functions across multiple different computing nodes through Java, whereas CUDA code helps accelerate local computations further on attached GPUs. All available heterogeneous computational power will be utilized. MapReduce in Hadoop eases the programming task by hiding communication and scheduling details. Hadoop Distributed File System will help achieve data-level fault resilience. GPU’s energy efficiency characteristics help reduce the power consumption of the whole system. To utilize GPU in Hadoop, four approaches including Jcuda, JNI, Hadoop Streaming, and Hadoop Pipes, have been accomplished and analyzed. Experimental results have demonstrated and compared their effectiveness.

Related Organizations

Chengdu University
China (People's Republic of)
Arkansas State University
United States
Providence University
Taiwan
Providence College
United States
Kansas State University
United States

View all View all

Keywords

Hadoop, Electronic computers. Computer science, GPU, MapReduce, CUDA, QA75.5-76.95

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	5
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average