
arXiv: 1903.12090
handle: 20.500.14243/359350
In information retrieval (IR) and related tasks, term weighting approaches typically consider the frequency of the term in the document and in the collection in order to compute a score reflecting the importance of the term for the document. In tasks characterized by the presence of training data (such as text classification) it seems logical that the term weighting function should take into account the distribution (as estimated from training data) of the term across the classes of interest. Although `supervised term weighting' approaches that use this intuition have been described before, they have failed to show consistent improvements. In this article we analyse the possible reasons for this failure, and call consolidated assumptions into question. Following this criticism we propose a novel supervised term weighting approach that, instead of relying on any predefined formula, learns a term weighting function optimised on the training set of interest; we dub this approach \emph{Learning to Weight} (LTW). The experiments that we run on several well-known benchmarks, and using different learning methods, show that our method outperforms previous term weighting approaches in text classification.
To appear in IEEE Transactions on Knowledge and Data Engineering
FOS: Computer and information sciences, Computer Science - Machine Learning, Term weighting, Deep learning, Machine Learning (stat.ML), Supervised term weighting, Computer Science - Information Retrieval, Machine Learning (cs.LG), Statistics - Machine Learning, Text classification, Neural networks, Information Retrieval (cs.IR)
FOS: Computer and information sciences, Computer Science - Machine Learning, Term weighting, Deep learning, Machine Learning (stat.ML), Supervised term weighting, Computer Science - Information Retrieval, Machine Learning (cs.LG), Statistics - Machine Learning, Text classification, Neural networks, Information Retrieval (cs.IR)
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 21 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Top 10% | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
