Views provided by UsageCounts
A set of 300 most frequent nouns has been extracted from the Russian National Corpus. Then, each method or resource, including RuThes and RuWordNet, produced at most five hypernyms, if possible. In case it is not possible, missing answers treated as empty results. This resulted in 10,600 unique non-empty subsumption pairs that have been passed for crowdsourcing annotation on the Yandex.Toloka microtask platform. Each pair has been annotated by seven different annotators whose mother tongue is Russian and the age is at least 20 by February 1, 2017. The layout of the human intelligence task (HIT) design assumes the direct answer to a simple question: does the given pair of words represent a meaningful is-a relation? Since the crowd workers are not expert lexicographers and this question might be difficult for them, it has been rephrased as “Is it correct that a kitten is a kind of mammal?” (in Russian). The answers have been aggregated using the Yandex.Toloka proprietary answer aggregation mechanism. As the result, 4,576 out of 10,600 pairs have been annotated as positive while the rest 6,024 have been annotated as negative. Interestingly, the workers were more confident in negative answers rather than in the positive ones. These negative answers are extremely useful for both training and testing different relation extraction methods. To the best of our knowledge, this is the first dataset of this kind made for the Russian language using microtask-based crowdsourcing.
Ustalov, D.: Expanding Hierarchical Contexts for Constructing a Semantic Word Network. In: Computational Linguistics and Intellectual Technologies: Papers from the Annual conference "Dialogue". Volume 1 of 2. Computational Linguistics: Practical Applications. pp. 369–381. RSUH, Moscow, Russia (2017)
is-a, Russian, crowdsourcing, natural language processing, hyponym, lexical relation, hypernym
is-a, Russian, crowdsourcing, natural language processing, hyponym, lexical relation, hypernym
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 11 |

Views provided by UsageCounts