Much Ado About Time: Exhaustive Annotation of Temporal Data
Gunnar A. Sigurdsson1, Olga Russakovsky1, Ali Farhadi2,3, Ivan Laptev4, Abhinav Gupta1,3 1Carnegie Mellon University 2University of Washington 3The Allen Institute for AI 4INRIA
Abstract
Large-scale annotated datasets allow AI systems to learn from and build upon the knowledge of the crowd. Many crowdsourcing techniques have been developed for collecting image annotations. These techniques often implicitly rely on the fact that a new input image takes a negligible amount of time to perceive. In contrast, we investigate and determine the most cost-effective way of obtaining high-quality multi-label annotations for temporal data such as videos. Watching even a short 30-second video clip requires a significant time investment from a crowd worker; thus, requesting multiple annotations following a single viewing is an important cost-saving strategy. But how many questions should we ask per video? We conclude that the optimal strategy is to ask as many questions as possible in a HIT (up to 52 binary questions after watching a 30-second video clip in our experiments). We demonstrate that while workers may not correctly answer all questions, the cost-benefit analysis nevertheless favors consensus from multiple such cheap-yet-imperfect iterations over more complex alternatives. When compared with a one-question-per-video baseline, our method is able to achieve a $10\%$ impr
中文速览
为大规模视频数据集收集多标签标注既耗时又昂贵——每次看完视频都要重新花时间,如果每次只问一个问题,成本极高。研究者提出了一种"一次看完视频、同时回答尽可能多问题"的众包标注策略,在30秒视频实验中最多同时提问52个二选一问题,并通过多名工人的结果取并集来弥补单个工人因认知负荷增加而漏答的问题。实验表明,与每次只问一个问题的基准方法相比,该策略在精确率相当(83.8% vs 83.0%)的情况下,召回率提升约10%(76.7% vs 66.7%),标注时间缩短近一半(3.8分钟 vs 7.1分钟),并成功用于对1815段视频完成了157类人体动作的完整多标签标注。这项工作为视频、音频、文本等需要投入较长感知时间的媒体数据的高效大规模标注提供了切实可行的方法论指导。
原文 arXiv:1607.07429;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1607.07429v2