What Question Answering can Learn from Trivia Nerds
iSchool, cs, umiacs, lsc University of Maryland、Jordan Boyd-Graber, ††\dagger Benjamin Börschinger† …………{jbg,、††\dagger Google Research, Zürich…………
Abstract
In addition to machines answering questions, question answering (qa) research creates interesting, challenging questions that reveal the best systems. We argue that creating a qa dataset—and its ubiquitous leaderboard—closely resembles running a trivia tournament: you write questions, have agents—humans or machines—answer questions, and declare a winner. However, the research community has ignored the lessons from decades of the trivia community creating vibrant, fair, and effective qa competitions. After detailing problems with existing qa datasets, we outline several lessons that transfer to qa research: removing ambiguity, discriminating skill, and adjudicating disputes.
中文速览
问答系统(Question Answering)研究领域热衷于建榜单、刷排名,却长期忽视了一个已经深耕数十年的"同行"——益智竞答(trivia tournament)社区。作者指出,学术界构建问答数据集、让机器或人类作答、公布排名的全套流程,本质上就是在办一场答题竞赛,因此可以直接借鉴竞答社区总结出的成熟经验:通过"试玩"(playtest)发现数据集中的捷径和歧义,设计能真正区分强弱选手的"金发姑娘难度"(Goldilocks)题目,以及用多维度指标而非单一分数来衡量系统能力。研究进一步引入"有效数据集比例"(ρ)这一概念,量化地说明标注错误与题目难度分布如何共同削弱榜单的鉴别力,并推荐了竞答黄金格式"学术碗"(Quizbowl)中的"金字塔式"(pyramidal)提问结构作为改进方案。这项工作的价值在于为问答研究提供了一套可操作的质量准则,帮助社区少走弯路,让榜单真正衡量智能而非测量数据噪声。
原文 arXiv:1910.14464;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1910.14464v3