Weighted shingling: an adaptation of shingling for weighted shingles

Zahra Eskandari Gharghe; Behrouz Minaei Bidgoli

Found an issue? Give us feedback

https://doi.org/10.1...arrow_drop_down

https://doi.org/10.1109/iit.20...

Article . 2009 . Peer-reviewed

Data sources: Crossref

https://dx.doi.org/10.1109/iit...

Article

Data sources: Microsoft Academic Graph

Weighted shingling: an adaptation of shingling for weighted shingles

descriptionPublicationkeyboard_double_arrow_right Article 01 Dec 2009Publisher:IEEEJournal:2009 International Conference on Innovations in Information Technology (IIT)

Authors: Zahra Eskandari Gharghe; Behrouz Minaei Bidgoli;

doi: 10.1109/iit.2009.5413370

Weighted shingling: an adaptation of shingling for weighted shingles

- Summary
- Metrics

Abstract

Broder's shingling is one of the state-of-the-art approaches in detecting near-duplicate documents. Prior evaluations of this method have shown that document-pairs which have different main content but have a large amount of similar unimportant details are the main sources of its errors. Different web pages from the same site are a good example of such documents. In such pages, almost always there is a similar boilerplate text which has a chance to be selected as the document's fingerprint and trick the algorithm. It seems that this problem is due to representing each document only by a sample of its shingles. This sample only contains some of the page's shingles and discards any other information. by Including additional information such as frequencies of shingles in this sample, we can improve the performance of the algorithm. This paper proposes a weighting of shingles and adapts shingling to be applied on weighted shingles. Our results have shown an improvement in shingling's performance.

Related Organizations

Iran University of Science and Technology
Iran (Islamic Republic of)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	1
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

1

Average

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

Upload OA version

Are you the author of this publication? Upload your Open Access version to Zenodo!

It’s fast and easy, just two clicks!

uploadUpload now