descriptionPublicationkeyboard_double_arrow_right Article , Preprint , Conference object 12 Apr 2024Embargo end date: 01 Jan 2023Publisher:ACMJournal:Proceedings of the IEEE/ACM 46th International Conference on Software Engineering

Authors: Youcef Remil; Anes Bendimerad; Romain Mathonat; Chedy Raïssi; Mehdi Kaytoue;

doi: 10.1145/3597503.3639146 , 10.48550/arxiv.2310.06703

arXiv: http://arxiv.org/abs/2310.06703

DeepLSH: Deep Locality-Sensitive Hash Learning for Fast and Efficient Near-Duplicate Crash Report Detection

- Summary
- Subjects
- Related research
  (1)
- Metrics

Abstract

Automatic crash bucketing is a crucial phase in the software development process for efficiently triaging bug reports. It generally consists in grouping similar reports through clustering techniques. However, with real-time streaming bug collection, systems are needed to quickly answer the question: What are the most similar bugs to a new one?, that is, efficiently find near-duplicates. It is thus natural to consider nearest neighbors search to tackle this problem and especially the well-known locality-sensitive hashing (LSH) to deal with large datasets due to its sublinear performance and theoretical guarantees on the similarity search accuracy. Surprisingly, LSH has not been considered in the crash bucketing literature. It is indeed not trivial to derive hash functions that satisfy the so-called locality-sensitive property for the most advanced crash bucketing metrics. Consequently, we study in this paper how to leverage LSH for this task. To be able to consider the most relevant metrics used in the literature, we introduce DeepLSH, a Siamese DNN architecture with an original loss function, that perfectly approximates the locality-sensitivity property even for Jaccard and Cosine metrics for which exact LSH solutions exist. We support this claim with a series of experiments on an original dataset, which we make available.

Related Organizations

French National Centre for Scientific Research
France
École Centrale de Lyon
France
Claude Bernard University Lyon 1
France
University of Lyon System
France
UNIVERSITE LUMIERE LYON 2
France

View all View all

Keywords

Locality-sensitive hashing, [INFO.INFO-AI] Computer Science [cs]/Artificial Intelligence [cs.AI], FOS: Computer and information sciences, Siamese neural networks, Approximate nearest neighbors, Computer Science - Artificial Intelligence, [INFO.INFO-SE] Computer Science [cs]/Software Engineering [cs.SE], [INFO] Computer Science [cs], Stack trace similarity, Software Engineering (cs.SE), Computer Science - Software Engineering, Artificial Intelligence (cs.AI), [INFO.INFO-CC] Computer Science [cs]/Computational Complexity [cs.CC], Crash deduplication Stack trace similarity Approximate nearest neighbors Locality-sensitive hashing Siamese neural networks, Crash deduplication

1 Research products, page 1 of 1

deep-locality-sensitive-hashing software on GitHub
IsRelatedTo

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	3
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average