
In this paper we describe a new approach to text categorization, our focus is in the amount of information (the entropy) in the text. The entropy is computed with the empirical distribution of words in the text. We provide the system with a manually segmented collection of documents in different categories. For each category a separate empirical distribution of words is computed, we will use this empirical distributions for categorization purposes. If we compute the entropy of the test document for each empirical distribution the correct category will show as a maximum. For example, if we compute the entropy of a sports document using the politics or the sports empirical word distributions then the computed entropy will be higher in sports than in politics. Our text categorization approach is simple, easy to code and needs no training time (aside from histogram computations). The classification time is linear on the size of the document and the number of document categories. We support our claims with extensive experimentation.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 3 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
