
Deduplication is a fundamental problem in data integration and record linkage, where the goal is to identify records that refer to the same real-world entity despite differences in formatting, wording, or completeness. This paper studies how threshold choice and auxiliary feature weighting affect match quality in tabular deduplication. Using the Abt-Buy benchmark dataset, we evaluate threshold-based matching with name similarity alone and compare it against weighted combinations of name, description, and price similarity. Our results show that threshold selection has a substantial impact on precision, recall, and F1, with poorly chosen thresholds leading to either excessive false matches or missed duplicates. We also find that additional features do not automatically improve performance. A naive multi-feature hybrid underperformed the name-only baseline, while a tuned name-plus-price model achieved the best results by improving recall with almost no loss of precision. A simple logistic regression comparison also underperformed the tuned threshold model, but its learned coefficients reinforced the same feature-level conclusion by assigning positive weight to name and price similarity and negative weight to description similarity. Together, these results show that deduplication quality depends not only on threshold choice, but also on whether added features provide reliable information beyond the primary matching field.
entity matching, deduplication, entity resolution, threshold tuning, record linkage, feature weighting
entity matching, deduplication, entity resolution, threshold tuning, record linkage, feature weighting
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
