Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
versions View all 2 versions
addClaim

Finding Reliable Sources: Evaluation Datasets

Authors: Przybyła, Piotr; Borkowski, Piotr; Kaczyński, Konrad;

Finding Reliable Sources: Evaluation Datasets

Abstract

Finding Reliable Sources (FRS) is an NLP/ML task, in which, given a single-sentence claim, we aim to automatically find reliable sources that could be used to confirm or refute it. This repository includes evaluation datasets that could be used to train and test solution for this problem. Datasets are organised as large collections of individual records, each consisting of (1) a claim expressed a short text fragment and (2) a list of identifiers (DOI, ISBN, arXiv ID or URL) of reliable sources associated with it. Two datasets are included: Claim-Source Pairing Dataset (CSP) contains 32 million textual contexts paired with 24 million source identifiers mined from English Wikipedia, FEVER-FRS contains 15,798 claims gathered for the FEVER shared task, supplemented with source identifiers using Wikipedia. The data were gathered within a study described in the article "Countering Disinformation by Finding Reliable Sources: a Citation-Based Approach", presented at the 2022 International Joint Conference on Neural Networks (IJCNN 2022). Please refer to the article (in conference proceedings or authors' version) for more information on the FRS task, data collection procedures and evaluation results. The research was done within the HOMADOS project at the Institute of Computer Science, Polish Academy of Sciences. The datasets described here are based on the Wikipedia Complete Citation Corpus (WCCC), which also includes the train/test split for CSP. The CSP dataset is uploaded in three variants, using different text fragments as 'context'. In 's', the context consists of a single sentence preceding (or containing) a citation. The 'ts' variant additionally includes a title of the Wikipedia article, while 'ss' contains two sentences preceding a citation. Each variant contains subsequently numbered TSV files, including lines with the following values: document ID (see WCCC for full document metadata), textual context, a list of (at least one, possibly many) source identifiers. The FEVER-FRS dataset was created based on the FEVER shared task data. In the original dataset, each claim is matched with a set of 'evidence' sentences from Wikipedia that either support or refute it. For the purpose of evaluating FRS solutions, we replaced each evidence sentence with the identifiers to external sources that are cited alongside it. This was not always possible (e.g. for Wikipedia sentences with no sources, see the paper for details), but we managed to find 4414 and 11384 claims that are, respectively, refuted or supported by the provided sources. Each line of the FEVER-FRS TSV files contains the following: text of a short claim, a list of (at least one, possibly many) source identifiers. The source code code for generating both datasets (CSP from Wikipedia and FEVER-FRS from FEVER) can be obtained from a GitHub repository.

Related Organizations
  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
    OpenAIRE UsageCounts
    Usage byUsageCounts
    visibility views 29
    download downloads 12
  • 29
    views
    12
    downloads
    Powered byOpenAIRE UsageCounts
Powered by OpenAIRE graph
Found an issue? Give us feedback
visibility
download
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
views
OpenAIRE UsageCountsViews provided by UsageCounts
downloads
OpenAIRE UsageCountsDownloads provided by UsageCounts
0
Average
Average
Average
29
12