
Content-based spam filtering technologies generally use feature selection algorithm for mail classification. Based on the mutual information feature selection algorithm, this paper proposes an improved mutual information method with frequency (MIf) by introducing the word frequency factor, and an improved mutual information method with average frequency (MIaf) by introducing the word average frequency factor. Simulation experiments are conducted based on the English corpus (PU1's lemm_stop) and Chinese corpus CCERT email data set, the feature subsets are extracted through the improved algorithms, and the mails are classified by the Naive Bayes algorithm. The experimental results show that the improved mutual information algorithms can select better feature subsets and enhance the mail classification effects.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 7 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
