Processing theta-joins using MapReduce

descriptionPublicationkeyboard_double_arrow_right Article , Conference object 12 Jun 2011Publisher:ACMJournal:Proceedings of the 2011 ACM SIGMOD International Conference on Management of data

Authors: Alper Okcan; Mirek Riedewald;

doi: 10.1145/1989323.1989423

Processing theta-joins using MapReduce

- Summary
- Related research
  (8)
- Metrics

Abstract

Joins are essential for many data analysis tasks, but are not supported directly by the MapReduce paradigm. While there has been progress on equi-joins, implementation of join algorithms in MapReduce in general is not sufficiently understood. We study the problem of how to map arbitrary join conditions to Map and Reduce functions, i.e., a parallel infrastructure that controls data flow based on key-equality only. Our proposed join model simplifies creation of and reasoning about joins in MapReduce. Using this model, we derive a surprisingly simple randomized algorithm, called 1-Bucket-Theta, for implementing arbitrary joins (theta-joins) in a single MapReduce job. This algorithm only requires minimal statistics (input cardinality) and we provide evidence that for a variety of join problems, it is either close to optimal or the best possible option. For some of the problems where 1-Bucket-Theta is not the best choice, we show how to achieve better performance by exploiting additional input statistics. All algorithms can be made 'memory-aware', and they do not require any modifications to the MapReduce environment. Experiments show the effectiveness of our approach.

Related Organizations

Northwestern University
United States
Northeastern University
United States

8 Research products, page 1 of 1

Toward fast theta‐join: A prefiltering and amalgamated partitioning approach
2022IsAmongTopNSimilarDocuments
An Efficient Theta-Join Query Processing Algorithm on MapReduce Framework
2012IsAmongTopNSimilarDocuments
Real-time business intelligence through compact and efficient query processing under updates.
2018IsAmongTopNSimilarDocuments
Conjunctive Queries with Theta Joins Under Updates
2019IsAmongTopNSimilarDocuments
Optimized Theta-Join Processing
2021IsAmongTopNSimilarDocuments
Efficient multi-way theta-join processing using MapReduce
2012IsAmongTopNSimilarDocuments
D-JB: An Online Join Method for Skewed and Varied Data Streams
2018IsAmongTopNSimilarDocuments
The $$\theta $$-Join as a Join with $$\theta $$
2020IsAmongTopNSimilarDocuments

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	171
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 1%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 1%

Found an issue? Give us feedback

171

Top 10%

Top 1%

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

Upload OA version

Are you the author of this publication? Upload your Open Access version to Zenodo!

It’s fast and easy, just two clicks!

uploadUpload now