
doi: 10.1109/cbd.2015.41
Since preserving data locality is very important in MapReduce frameworks, the placement of jobs' input data becomes critical for MapReduce job performance. However, existing data placement strategies applied in MapReduce based file systems lack the efficient control of data storage. Some input data of jobs may concentrate in a few nodes, making slots contention on those nodes more aggressive and thereby hurting job performance significantly. To address this problem, we propose a Slot-Sensitive Data Placement strategy that aims to alleviate the contention for slots due to inappropriate data placement. By separately placing the input data of jobs on as many nodes as possible according to the availability of slots, SSDP contributes to the improvement of job performance. Extensive simulations by replaying public workloads demonstrate that SSDP can significantly improve the job performance up to 13%.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
