Examining Sentiment in Complex Texts. A Comparison of Different Computational Approaches

descriptionPublicationkeyboard_double_arrow_right Article , Conference object , Other literature type 04 May 2022Publisher:Frontiers Media SAJournal:Frontiers in Big Data, volume 5 (eissn: 2624-909X,

Copyright policy )

Authors: Stefan Munnes; Corinna Harsch; Marcel Knobloch; Johannes S. Vogel; Johannes S. Vogel; Lena Hipp; Lena Hipp; +1 Authors

APC: 1,261.17 EUR

doi: 10.3389/fdata.2022.886362

pmid: 35600329

pmc: PMC9114298

handle: 10419/261091

Examining Sentiment in Complex Texts. A Comparison of Different Computational Approaches

- Summary
- Subjects
- Metrics

Abstract

Can we rely on computational methods to accurately analyze complex texts? To answer this question, we compared different dictionary and scaling methods used in predicting the sentiment of German literature reviews to the “gold standard” of human-coded sentiments. Literature reviews constitute a challenging text corpus for computational analysis as they not only contain different text levels—for example, a summary of the work and the reviewer's appraisal—but are also characterized by subtle and ambiguous language elements. To take the nuanced sentiments of literature reviews into account, we worked with a metric rather than a dichotomous scale for sentiment analysis. The results of our analyses show that the predicted sentiments of prefabricated dictionaries, which are computationally efficient and require minimal adaption, have a low to medium correlation with the human-coded sentiments (r between 0.32 and 0.39). The accuracy of self-created dictionaries using word embeddings (both pre-trained and self-trained) was considerably lower (r between 0.10 and 0.28). Given the high coding intensity and contingency on seed selection as well as the degree of data pre-processing of word embeddings that we found with our data, we would not recommend them for complex texts without further adaptation. While fully automated approaches appear not to work in accurately predicting text sentiments with complex texts such as ours, we found relatively high correlations with a semiautomated approach (r of around 0.6)—which, however, requires intensive human coding efforts for the training dataset. In addition to illustrating the benefits and limits of computational approaches in analyzing complex text corpora and the potential of metric rather than binary scales of text sentiment, we also provide a practical guide for researchers to select an appropriate method and degree of pre-processing when working with complex texts.

Related Organizations

University of Erlangen-Nuremberg
Germany
Friedrich-Alexander-Universität Erlangen-Nürnberg
Social Science Research Center Berlin
Germany
University of Potsdam
Germany
Leibniz Association
Germany

View all View all

Keywords

word embeddings, Big Data, ddc:300, automated text analysis, Information technology, German literature, T58.5-58.64, computer-assisted text analysis, scaling method, sentiment analysis, Sozialwissenschaften, dictionary

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	7
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 10%

Found an issue? Give us feedback

7

Top 10%

Average

Top 10%

Green

gold

Fields of Science (3) View all

Fields of Science