
handle: 11573/498840
Errors can be detected into large data sets by means of rule-based techniques (e.g. Fellegi-Holt). They essentially consist in checking if each record satisfies a number of rules. Records not respecting all the rules are declared erroneous records. In all the cases When data collecting has a cost, which are the majority, we are interested in correcting such data. The correction consists in changing a number of fields of the erroneous record, until it satisfies the above rules. This should generally be performed by modifying as less as possible the erroneous data, while causing minimum perturbation to the original frequency distributions of the data. Such process is called data imputation, and is of great relevance in the field of statistics. A new procedure for data imputation by using a discrete optimization model is here presented.
Information Reconstruction; Combinatorial Optimization; Data Mining
Information Reconstruction; Combinatorial Optimization; Data Mining
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
