Efficient unsupervised discovery of word categories using symmetric patterns and high frequency words

descriptionPublicationkeyboard_double_arrow_right Article , Conference object 01 Jan 2006Publisher:Association for Computational Linguistics (ACL)Journal:Proceedings of the 21st International Conference on Computational Linguistics and the 44th annual meeting of the ACL - ACL '06

Authors: Dmitry Davidov; Ari Rappoport;

doi: 10.3115/1220175.1220213

Efficient unsupervised discovery of word categories using symmetric patterns and high frequency words

- Summary
- Metrics

Abstract

We present a novel approach for discovering word categories, sets of words sharing a significant aspect of their meaning. We utilize meta-patterns of high-frequency words and content words in order to discover pattern candidates. Symmetric patterns are then identified using graph-based measures, and word categories are created based on graph clique sets. Our method is the first pattern-based method that requires no corpus annotation or manually provided seed patterns or words. We evaluate our algorithm on very large corpora in two languages, using both human judgments and WordNet-based evaluation. Our fully unsupervised results are superior to previous work that used a POS tagged corpus, and computation time for huge corpora are orders of magnitude faster than previously reported.

Related Organizations

Hebrew University of Jerusalem
Israel

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	24
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

24

Average

Top 10%

Average

bronze

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering