
handle: 1842/38843
Newspaper comments are content generated and posted by users in response to a news- paper article. The volume of comments posted has increased to the point where they are not easily readable or understandable within the systems currently used. The re- search that is conducted here is done with the aim of summarising this valuable content in order to retain the value, usefulness and usability of the comments as their number increases. The amount of research within the newspaper comment summarisation domain has increased in the last three years, including the creation of a shared task focused in this area. In this thesis we share our contribution to this effort. We describe the newspaper summarisation task as having three distinct phases: 1) the identification of topics discussed within the comments, 2) a description of the topics which serves as a summary, and 3) an evaluation of how well the task has been conducted. We initially analyse the approaches taken and the gold standard data from the shared task, OnForumS. This shared task asked researchers to 1) Link comment sen- tences to article sentences or other comment sentences in order that the linked sen- tences form topics, 2) Type the links that are found by labelling them with both ar- gument type (in favour or against) and sentiment (positive or negative). The second phase would likely aid in the description of the topics. A selection of links and link types were validated using crowd sourcing and the crowd source data was offered to the community as a gold standard for future evaluation. We found that whilst we were able to reimplement the most successful approaches taken in this task we obtained a different pattern of results. We found that the type of validation performed did not give the full picture of how well the approaches performed and the further evaluation we conducted raised questions about the suitability of both the structure of the task and the data provided for evaluation. We concluded that this was a very difficult task to evaluate fully and the data provided was not a suitable gold standard for use in the remainder of this work. When we compared clustering approaches for topic analysis of the comment data we found that LDA topic modelling was the most successful approach, when compared with K-means clustering, GAAC clustering, comments grouped on word frequency, and a random baseline. This was established via evaluation using five comment sets that had been annotated for topic. We investigated ways to enhance the LDA topic modelling and found that if the comments were aggregated into documents that merged comments with their direct replies this gave better topic models. We tackled the description of topic phase by extracting sentences from the com- ments to represent the topic clusters. We compared MI, TF-IDF, MMR, TextRank and a random baseline and found that no approach consistently performed well. We combined the three best performing approaches (MI, TF-IDF and TextRank) and this ensemble outperformed the other approaches. We evaluate both the topic identification subtask and the topic description subtask using intrinsic and extrinsic measures. In both tasks we analysed the results using mul- tiple metrics in order to evaluate both the approach and the suitability of the metric to judge the approach. We often found that the metrics gave different results but com- monly the ranked order of approaches was the same. As a consequence in future work we would use: for topic analysis, F-score with a mapping of one automatic cluster to one gold cluster where gold standard data is available and perplexity where not, for topic description we would use a combination of ROUGE and F-score to compare with a gold standard, and a human ranking judgement to evaluate the final summaries.
summarisation, user gee rated content, LDA, social media, newspaper comments, topic modelling
summarisation, user gee rated content, LDA, social media, newspaper comments, topic modelling
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
