Computing marginals using MapReduce

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 01 Jun 2018Embargo end date: 01 Jan 2015 English Publisher:Elsevier BVJournal:Journal of Computer and System Sciences, volume 94, pages 98-117 (issn: 0022-0000,

Copyright policy )

Authors: Foto N. Afrati; Shantanu Sharma 0001; Jonathan R. Ullman; Jeffrey D. Ullman;

doi: 10.1016/j.jcss.2017.02.007 , 10.48550/arxiv.1509.08855

arXiv: 1509.08855

Computing marginals using MapReduce

- Summary
- Subjects
- Metrics

Abstract

We consider the problem of computing the data-cube marginals of a fixed order $k$ (i.e., all marginals that aggregate over $k$ dimensions), using a single round of MapReduce. The focus is on the relationship between the reducer size (number of inputs allowed at a single reducer) and the replication rate (number of reducers to which an input is sent). We show that the replication rate is minimized when the reducers receive all the inputs necessary to compute one marginal of higher order. That observation lets us view the problem as one of covering sets of $k$ dimensions with sets of a larger size $m$, a problem that has been studied under the name "covering numbers." We offer a number of constructions that, for different values of $k$ and $m$ meet or come close to yielding the minimum possible replication rate for a given reducer size.

Related Organizations

Ben-Gurion University of the Negev
Israel
Northwestern University
United States
Northwestern University
United States
Northwestern University
University of California, Irvine
United States

View all View all

Keywords

FOS: Computer and information sciences, Computer Science - Databases, General topics in the theory of data, MapReduce, Databases (cs.DB), marginals, data-cube

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	4
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

4

Average

Green

bronze

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering