Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Other literature type . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Other literature type . 2026
License: CC BY
Data sources: Datacite
ZENODO
Other literature type . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

Evaluating the Effect of Deduplication Thresholds and Feature Weighting on Match Quality

Authors: Lin, Kelvin;

Evaluating the Effect of Deduplication Thresholds and Feature Weighting on Match Quality

Abstract

Deduplication is a fundamental problem in data integration and record linkage, where the goal is to identify records that refer to the same real-world entity despite differences in formatting, wording, or completeness. This paper studies how threshold choice and auxiliary feature weighting affect match quality in tabular deduplication. Using the Abt-Buy benchmark dataset, we evaluate threshold-based matching with name similarity alone and compare it against weighted combinations of name, description, and price similarity. Our results show that threshold selection has a substantial impact on precision, recall, and F1, with poorly chosen thresholds leading to either excessive false matches or missed duplicates. We also find that additional features do not automatically improve performance. A naive multi-feature hybrid underperformed the name-only baseline, while a tuned name-plus-price model achieved the best results by improving recall with almost no loss of precision. A simple logistic regression comparison also underperformed the tuned threshold model, but its learned coefficients reinforced the same feature-level conclusion by assigning positive weight to name and price similarity and negative weight to description similarity. Together, these results show that deduplication quality depends not only on threshold choice, but also on whether added features provide reliable information beyond the primary matching field.

Keywords

entity matching, deduplication, entity resolution, threshold tuning, record linkage, feature weighting

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green