Constructing Datasets for Multi-hop Reading Comprehension Across Documents

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 01 Dec 2018Embargo end date: 01 Jan 2017 United Kingdom English Publisher:MIT Press - JournalsJournal:Transactions of the Association for Computational Linguistics, volume 6, pages 287-302 (eissn: 2307-387X,

Copyright policy )Funded by:EC | SUMMA

Authors: Welbl, J; Saito Stenetorp, PLEPS; Riedel, S;

doi: 10.1162/tacl_a_00021 , 10.48550/arxiv.1710.06481

arXiv: 1710.06481

Constructing Datasets for Multi-hop Reading Comprehension Across Documents

- Summary
- Subjects
- Metrics

Abstract

Most Reading Comprehension methods limit themselves to queries which can be answered using a single sentence, paragraph, or document. Enabling models to combine disjoint pieces of textual evidence would extend the scope of machine comprehension methods, but currently no resources exist to train and test this capability. We propose a novel task to encourage the development of models for text understanding across multiple documents and to investigate the limits of existing methods. In our task, a model learns to seek and combine evidence — effectively performing multihop, alias multi-step, inference. We devise a methodology to produce datasets for this task, given a collection of query-answer pairs and thematically linked documents. Two datasets from different domains are induced, and we identify potential pitfalls and devise circumvention strategies. We evaluate two previously proposed competitive models and find that one can integrate information across documents. However, both models struggle to select relevant information; and providing documents guaranteed to be relevant greatly improves their performance. While the models outperform several strong baselines, their best accuracy reaches 54.5% on an annotated test set, compared to human performance at 85.0%, leaving ample room for improvement.

Country

United Kingdom

Related Organizations

University College London
United Kingdom
UNIVERSITY COLLEGE LONDON, Bartlett School of Planning
United Kingdom

Keywords

FOS: Computer and information sciences, Computer Science - Computation and Language, Artificial Intelligence (cs.AI), Computer Science - Artificial Intelligence, Computational linguistics. Natural language processing, P98-98.5, Computation and Language (cs.CL)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	188
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 0.1%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 0.1%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 1%

Found an issue? Give us feedback

188

Top 0.1%

Top 1%

Green

gold

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

Fields of Science

engineering and technology

electrical engineering, electronic engineering, information engineering

Funded by

EC| SUMMA