
This paper proposed a method for building domain-specific thesauri automatically from plain text corpus based on Latent Dirichlet Allocation (LDA). This method consists of two steps: 1) discovering domain-specific terms from document collections of multiple domains, and 2) learning hierarchical relations between the associated terms of each domain. The novelty of step 1 lies in the utilization of LDA in selecting terms with high predictive probability of a specific domain via latent topics, which overcomes the drawbacks of unigram model. Meanwhile, the hierarchical relations among domain terms are exploited by a novel approach based on word association analysis in step 2. The proposed method is tested on two datasets in different languages. The experimental results show that the terms obtained by this method are intuitively relevant to the reference domain and many term pairs with hierarchical relations are discovered. And the relations reflect the structure of the domain rather well. Compared to other approaches, the proposed one is more accurate in both domain terms mining and hierarchical relation learning tasks.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 1 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
