Improving query expansion using WordNet

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 13 May 2014Embargo end date: 01 Jan 2013 English Publisher:WileyJournal:Journal of the Association for Information Science and Technology, volume 65, pages 2,469-2,478 (issn: 2330-1635, eissn: 2330-1643,

Copyright policy )

Authors: Dipasree Pal; Mandar Mitra; Kalyankumar Datta;

doi: 10.1002/asi.23143 , 10.48550/arxiv.1309.4938

arXiv: 1309.4938

Improving query expansion using WordNet

- Summary
- Subjects
- Metrics

Abstract

This study proposes a new way of using WordNet for query expansion (QE). We choose candidate expansion terms from a set of pseudo‐relevant documents; however, the usefulness of these terms is measured based on their definitions provided in a hand‐crafted lexical resource such as WordNet. Experiments with a number of standard TREC collections WordNet‐based that this method outperforms existing WordNet‐based methods. It also compares favorably with established QE methods such as KLD and RM3. Leveraging earlier work in which a combination of QE methods was found to outperform each individual method (as well as other well‐known QE methods), we next propose a combination‐based QE method that takes into account three different aspects of a candidate expansion term's usefulness: (a) its distribution in the pseudo‐relevant documents and in the target corpus, (b) its statistical association with query terms, and (c) its semantic relation with the query, as determined by the overlap between the WordNet definitions of the term and query terms. This combination of diverse sources of information appears to work well on a number of test collections, viz., TREC123, TREC5, TREC678, TREC robust (new), and TREC910 collections, and yields significant improvements over competing methods on most of these collections.

Related Organizations

Indian Statistical Institute
India
Jadavpur University
India

Keywords

FOS: Computer and information sciences, Information Retrieval (cs.IR), Computer Science - Information Retrieval

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	46
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%