
Question Understanding of Chinese Question-Answering System generally includes steps such as: word segmentation, POS Tagging, keywords expansion, information retrieval etc. The extended keyword set usually has redundant messages and part of the words and phrases may be not relevant to the question. Consequently, information retrieval with the extended keywords set may bring about large numbers of noise information and enhance the difficulty of answer pick-up. This paper explores the use of distance between vocabularies in the sememe tree for reducing keywords set. It analyzes the detailed steps of question understanding and the improved algorithm. Empirical results support the theoretical findings. The algorithm proposed in the paper achieves substantial improvement by 23% on the average, and wipes off the vocabulary beside the mark. Furthermore, it will improve the accuracy rate of Question Understanding in the subsequent steps.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
