
Missing value is a widespread problem for data quality because most of the statistical procedures require a value for each variable. The missing value may lead to biased parameter estimates, and may result in degradation of data quality. Imputation has been used to replace the missing data by plausible estimation. This research combines the Hot-Deck and Expectation Maximization imputations with the C5.0 classifier technique to estimate missing values and to improve the data quality. It would fit numeric, categorical, and continuous data sets. The Hot-Deck imputation deals with categorical and mixed data types. The Expectation Maximization imputation is a best method for numerical data and to increase association with other variables. The C5.0 classifies the data in lesser time with minimum memory usage. It has a higher accuracy compared to other classifiers. This new embedded ‘Intelligent Imputation Technique for Missing Values’ is used for the main process of acquiring knowledge from data. This technique is compared with the original C5.0 algorithm and the results are presented.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 2 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
