Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ Harvard Dataversearrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
Harvard Dataverse
Dataset . 2022
License: CC 0
Data sources: Datacite
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Dataset . 2022
License: CC BY
Data sources: Datacite
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Dataset . 2022
License: CC BY
Data sources: Datacite
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Dataset . 2022
License: CC BY
Data sources: ZENODO
versions View all 3 versions
addClaim

One million articles from five post socialist countries with extracted features: sentiment, basic emotions, LDA topics and presence of influential domestic politicians

Authors: blind review, Under;

One million articles from five post socialist countries with extracted features: sentiment, basic emotions, LDA topics and presence of influential domestic politicians

Abstract

This is a replication data for my paper under blind review. <br><br> Abstract This paper develops a new prediction model for media content presence on a website. It analyses a new corpus of one million articles from five countries: Poland, Russia, Belarus, Kazakhstan and Ukraine, in two languages, Polish and Russian. These articles were scraped daily from seventeen websites in 2017-2020 period. The research applies a wide range of natural language processing methods to automatically derive several properties of each article: its topic, sentiment, basic emotions, mentions of influential domestic politicians. The articles’ embeddings and their cosine similarity are used to calculate the news context, such as how an article differs from the daily issue main themes. These features are used to estimate a logistic regression assessing the likelihood that the same or slightly modified, as measured by cosine similarity, article will remain on the main web page the next day. The key, and somewhat unexpected result is that articles with negative sentiment polarity are less likely to be published for more than one day. This result holds for all countries analyzed. It means that the negative news bias documented in the literature is partly offset by their shorter life cycle.<br><br> Data is in the Python pickle format. Should be read into Python using the pickle.load() function. Each element (row) is the data frames or list represents one news article. Each file has the same format. Loading a pickle file returns a list of four elements:<br> 1. A dummy variable equal to 1 when the article was published the next day, with the text being identical<br> 2. A dummy variable equal to 1 when the article was published the next day, but we allow for small text modifications (cosine similarity > 0.99)<br> 3. Dataframe with extracted features, described below.<br> 4. List with texts of articles in Polish or Russian<br><br> Ad 3. The columns of the dataframe are as follows (we refer to row number i in description):<br> - pandas index (may appear once or twice in the datafame)<br> - maxcosine: maximum cosine similarity between art i and all articles published next day<br> - cosine_diff: cosine similarity between article i and the elementwise average of embeddings of all articles in the current issue. Measure how similar is the article i to the core narrative of the current issue<br> - cosine_std: std. dev. of cosine similarity measures between all pairs of articles in the current issue. Measures how focused or dispersed is the current issue news coverage<br> - thirteen LDA topic groups: politics, legislation and legal affairs (POL); economy, finance, various sectors of the economy (ECO); military, war, protests, crime, security threats (MIL); international affairs, specific issues concerning foreign countries (INT); technology (TECH); family issues, culture, sport, education (FAM); regional issues and housing (REG); health issues and the Covid-19 pandemic (HEA); media (MED); accidents (ACC); religion (REL); the Soviet Union (USSR); and articles for which no topic could be determined (MISC). <br> - rsent.c: relative sentiment that is dictionary based sentiment of articles i minus the average sentiment of the newspaper. This approach eliminates newspaper or country idiosyncratic sentiment factors. c stands for Covid, the sentiment lexicon was augmented with Covid related terms <br> - dip_*: Variable measuring if influential domestic politicians are mentioned in article i, * represent a country acronym. If N is equal to the number of occurrences of the names of influential domestic politicians in the article i, dip_* = 0 if N=0, dip_* = 1+ log(N) if N>0. <br> - three or four names of news portals from which the data was scraped.<br> - names of six basic emotions and the article i emotion scores calculated using zero-shot learning and the large version of the XLM (Conneau et al., 2019) model from the huggingface transformers library available at https://huggingface.co/vicgalle/xlm-roberta-large-xnli-anli <br> Names of the politicians used to calculate dip variables<br> Russia<br> "putin" "medvedev" "vaino" "shoigu" "bortnikov" "lavrov" "mishustin" "kirienko" "sechin" <br> Ukraine<br> "zelensky" "shmygal" "akhmetov" "avakov" "ermak" "poroshenko" "medvedchuk" "groisman" <br> Kazakhstan<br> "sagyntaev" "mamin" "tokayev" "nnazarbayev" "dnazarbayeva" "kulibayev" "masimov" <br> Belarus<br> "alukashenko" "vakulchik" "vlukashenko" "kobyakov" "makei" "myasnikovich" [37] "rumas" "golovchenko" <br> Poland<br> "kaczynski" "duda" "morawiecki" "ziobro" <br> Data coverage<br> Country, news portal, numbr of articles<br> Russia iz.ru 43,782<br> Russia kommersant.ru 46,070<br> Russia novayagazeta.ru 29,357<br> Russia vedomosti.ru 27,797<br> Kazakhstan informburo.kz 29,375<br> Kazakhstan nur.kz 67,350<br> Kazakhstan tengrinews.kz 44,285<br> Kazakhstan zakon.kz 109,442<br> Belarus bdg.by 33,447<br> Belarus belgazeta.by 21,995<br> Belarus sb.by 83,685<br> Ukraine kp.ua 194,792<br> Ukraine segodnya.ua 45,835<br> Ukraine vesti.ua 90,559<br> Poland gazeta.pl 53,321<br> Poland rp.pl 49,587<br> Poland wpolityce.pl 76,625<br><br> In the provided dataframes the number of observations is smaller, because the issues for which there was no next day issue, were removed. <br><br> Data was scraped daily between 2017 or 2018 (depending on the country) and January 2021.<br>

Keywords

Computer and Information Science, Social Sciences

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
    OpenAIRE UsageCounts
    Usage byUsageCounts
    visibility views 23
    download downloads 2
  • 23
    views
    2
    downloads
    Powered byOpenAIRE UsageCounts
Powered by OpenAIRE graph
Found an issue? Give us feedback
visibility
download
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
views
OpenAIRE UsageCountsViews provided by UsageCounts
downloads
OpenAIRE UsageCountsDownloads provided by UsageCounts
0
Average
Average
Average
23
2