Similarity Join and Similarity Self-Join Size Estimation in a Streaming Environment

descriptionPublicationkeyboard_double_arrow_right Article , Preprint , Other literature type 01 Apr 2020Embargo end date: 01 Jan 2018Publisher:Institute of Electrical and Electronics Engineers (IEEE)Journal:IEEE Transactions on Knowledge and Data Engineering, volume 32, pages 768-781 (issn: 1041-4347, eissn: 2326-3865,

Copyright policy )Funded by:NSERC | unidentified

Authors: Davood Rafiei; Fan Deng 0004;

doi: 10.1109/tkde.2019.2893175 , 10.48550/arxiv.1806.03313

arXiv: 1806.03313

Similarity Join and Similarity Self-Join Size Estimation in a Streaming Environment

- Summary
- Subjects
- Metrics

Abstract

We study the problem of similarity self-join and similarity join size estimation in a streaming setting where the goal is to estimate, in one scan of the input and with sublinear space in the input size, the number of record pairs that have a similarity within a given threshold. The problem has many applications in data cleaning and query plan generation, where the cost of a similarity join may be estimated before actually doing the join. On unary input where two records either match or don't match, the problem becomes join and self-join size estimation for which one-pass algorithms are readily available. Our work addresses the problem for d-ary input, for d >= 1, where the degree of similarity can vary from 1 to d. We show that our proposed algorithm gives an accurate estimate and scales well with the input size. We provide error bounds and time and space costs, and conduct an extensive experimental evaluation of our algorithm, comparing its estimation accuracy to a few competitors, including some multi-pass algorithms. Our results show that given the same space, the proposed algorithm has an order of magnitude less error for a large range of similarity thresholds.

IEEE Transactions on Knowledge and Data Engineering (to appear)

Related Organizations

University of Alberta
Canada

Keywords

FOS: Computer and information sciences, Computer Science - Databases, Databases (cs.DB)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	5
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%