
arXiv: 1407.7930
handle: 2434/495564 , 11568/790769 , 2158/1149242
Entity-linking is a natural-language-processing task that consists in identifying the entities mentioned in a piece of text, linking each to an appropriate item in some knowledge base; when the knowledge base is Wikipedia, the problem comes to be known as wikification (in this case, items are wikipedia articles). One instance of entity-linking can be formalized as an optimization problem on the underlying concept graph, where the quantity to be optimized is the average distance between chosen items. Inspired by this application, we define a new graph problem which is a natural variant of the Maximum Capacity Representative Set. We prove that our problem is NP-hard for general graphs; nonetheless, under some restrictive assumptions, it turns out to be solvable in linear time. For the general case, we propose two heuristics: one tries to enforce the above assumptions and another one is based on the notion of hitting distance; we show experimentally how these approaches perform with respect to some baselines on a real-world dataset.
In Proceedings GRAPHITE 2014, arXiv:1407.7671. The second and third authors were supported by the EU-FET grant NADINE (GA 288956)
FOS: Computer and information sciences, General Engineering, QA75.5-76.95, G.2.2, G.2.3, Computer Science - Information Retrieval, G.2.2; G.2.3; F.2.m, Electronic computers. Computer science, Computer Science - Data Structures and Algorithms, QA1-939, General Earth and Planetary Sciences, Data Structures and Algorithms (cs.DS), Software, Mathematics, F.2.m, Information Retrieval (cs.IR), General Environmental Science
FOS: Computer and information sciences, General Engineering, QA75.5-76.95, G.2.2, G.2.3, Computer Science - Information Retrieval, G.2.2; G.2.3; F.2.m, Electronic computers. Computer science, Computer Science - Data Structures and Algorithms, QA1-939, General Earth and Planetary Sciences, Data Structures and Algorithms (cs.DS), Software, Mathematics, F.2.m, Information Retrieval (cs.IR), General Environmental Science
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 1 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
