Lost in Transduction: Transductive Transfer Learning in Text Classification

descriptionPublicationkeyboard_double_arrow_right Article , Other literature type 20 Jul 2021 Italy English Publisher:Association for Computing Machinery (ACM)Journal:ACM Transactions on Knowledge Discovery from Data, volume 16, pages 1-21 (issn: 1556-4681, eissn: 1556-472X,

Copyright policy )Funded by:EC | ARIADNEplus, EC | AI4Media, EC | SoBigData-PlusPlus +1 projects

Authors: Moreo, Alejandro; Esuli, Andrea; Sebastiani, Fabrizio;

doi: 10.1145/3453146

handle: 20.500.14243/399577

Lost in Transduction: Transductive Transfer Learning in Text Classification

- Summary
- Subjects
- Metrics

Abstract

Obtaining high-quality labelled data for training a classifier in a new application domain is often costly. Transfer Learning (a.k.a. “Inductive Transfer”) tries to alleviate these costs by transferring, to the “target” domain of interest, knowledge available from a different “source” domain. In transfer learning the lack of labelled information from the target domain is compensated by the availability at training time of a set of unlabelled examples from the target distribution. Transductive Transfer Learning denotes the transfer learning setting in which the only set of target documents that we are interested in classifying is known and available at training time. Although this definition is indeed in line with Vapnik’s original definition of “transduction”, current terminology in the field is confused. In this article, we discuss how the term “transduction” has been misused in the transfer learning literature, and propose a clarification consistent with the original characterization of this term given by Vapnik. We go on to observe that the above terminology misuse has brought about misleading experimental comparisons, with inductive transfer learning methods that have been incorrectly compared with transductive transfer learning methods. We then, give empirical evidence that the difference in performance between the inductive version and the transductive version of a transfer learning method can indeed be statistically significant (i.e., that knowing at training time the only data one needs to classify indeed gives an advantage). Our clarification allows a reassessment of the field, and of the relative merits of the major, state-of-the-art algorithms for transfer learning in text classification.

Country

Italy

Related Organizations

Institute of Information Science and Technologies "A. Faedo"
Italy
National Research Council
Romania
National Research Council
Italy

Keywords

Transduction, Text classification, Distributional hypothesis, Induction, Transfer learning

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	13
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%