A Survey of NLP-Related Crowdsourcing HITs: what works and what does not
Jessica Huynh Carnegie Mellon University \AndJeffrey Bigham Carnegie Mellon University \AndMaxine Eskenazi Carnegie Mellon University
Abstract
Crowdsourcing requesters on Amazon Mechanical Turk (AMT) have raised questions about the reliability of the workers. The AMT workforce is very diverse and it is not possible to make blanket assumptions about them as a group. Some requesters now reject work en mass when they do not get the results they expect. This has the effect of giving each worker (good or bad) a lower Human Intelligence Task (HIT) approval score, which is unfair to the good workers. It also has the effect of giving the requester a bad reputation on the workers’ forums. Some of the issues causing the mass rejections stem from the requesters not taking the time to create a well-formed task with complete instructions and/or not paying a fair wage. To explore this assumption, this paper describes a study that looks at the crowdsourcing HITs on AMT that were available over a given span of time and records information about those HITs. This study also records information from a crowdsourcing forum on the worker perspective on both those HITs and on their corresponding requesters. Results reveal issues in worker payment and presentation issues such as missing instructions or HITs that are not doable.
中文速览
亚马逊众包平台(Amazon Mechanical Turk,AMT)上的任务发布者常常抱怨工人工作质量差,甚至批量拒绝付款,但这种做法对认真工作的工人极不公平。本文通过一周内系统记录AMT上的自然语言处理相关任务(HIT),并结合工人评价网站TurkerView的反馈,从发布者和工人两个视角分析问题根源。结果发现,30%的可查看任务存在技术故障或说明不清等问题,44%的任务时薪低于美国联邦最低工资标准7.25美元,还有26%的发布者与工人沟通评价较差。这说明数据质量低下的责任往往在于发布者——任务设计粗糙、薪酬不合理——而非工人本身,呼吁研究者在使用众包平台时认真设计任务、合理支付报酬,以保障数据质量和工人权益。
原文 arXiv:2111.05241;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2111.05241v1