To Re(label), or Not To Re(label)

Christopher H. Lin; Mausam; Daniel S. Weld

Found an issue? Give us feedback

Proceedings of the A...arrow_drop_down

Proceedings of the AAAI Conference on Human Computation and Crowdsourcing

Article . 2014 . Peer-reviewed

Data sources: Crossref

DBLP

Conference object

Data sources: DBLP

To Re(label), or Not To Re(label)

descriptionPublicationkeyboard_double_arrow_right Article , Conference object 05 Sep 2014Publisher:Association for the Advancement of Artificial Intelligence (AAAI)Journal:Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, volume 2, pages 151-158 (issn: 2769-1330, eissn: 2769-1349,

Copyright policy )Funded by:NSF | RI: Small: Integrating Pa..., NSF | RI: Small: Decision-Theor..., NSF | RI: Small: Improving Crow...

Authors: Christopher H. Lin; Mausam; Daniel S. Weld;

doi: 10.1609/hcomp.v2i1.13167

To Re(label), or Not To Re(label)

- Summary
- Metrics

Abstract

One of the most popular uses of crowdsourcing is to provide training data for supervised machine learning algorithms. Since human annotators often make errors, requesters commonly ask multiple workers to label each example. But is this strategy always the most cost effective use of crowdsourced workers? We argue "No" --- often classifiers can achieve higher accuracies when trained with noisy "unilabeled" data. However, in some cases relabeling is extremely important. We discuss three factors that may make relabeling an effective strategy: classifier expressiveness, worker accuracy, and budget.

Related Organizations

Washington State University
United States
University of Washington
United States
Washington State University Spokane
United States
Indian Institutes of Technology
India
Indian Institute of Technology Dharwad
India

View all View all

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	12
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average